In Reply: Can Artificial Intelligence Make the Cut? Dissecting Large Language Model’s Surgical Exam Performance
Adam M. Ostrovsky, Joshua R. Chen, Vishal N. Shah, Babak Abai
Abstract
To the Editor: We appreciate the thoughtful and comprehensive response to our paper, Performance of 5 prominent large language models in surgical knowledge evaluation: a comparative analysis.1 We are grateful for the opportunity to address the important points raised by the authors of the letter.2 We concur with the observation that the variability in responses from different large language models (LLMs) to identical queries poses a substantial reliability concern. As highlighted in our study, the inconsistency in responses across multiple trials indeed challenges the use of these models as de
§ The Valyu brief
Reading the full paper and taking notes. This takes a few seconds…
§ Ask this paper
Ask a question about this paper
Valyu reads the full text and answers from what the paper actually says.
Searching the other archives…