Paraphernalia
AarXiv15 Feb 2024Cited 6×

Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4

Ting Fang Tan, Kabilan Elangovan, Liyuan Jin, Jie Yao, Yong Li, Joshua Lim, Stanley Poh, Wei Yan Ng, Daniel V. Lim, Yuhe Ke, Nan Liu, Daniel Shu Wei Ting

Abstract

in Ophthalmology and LLM-based evaluation using GPT-4 Authors: ['Ting Fang Tan' 'Kabilan Elangovan' 'Liyuan Jin' 'Jie Yao' 'Yong Li' 'Joshua Lim' 'Stanley Poh' 'Wei Yan Ng' 'Daniel V. Lim' 'Yuhe Ke' 'Nan Liu' 'Daniel Shu Wei Ting'] Methods: A dataset of 400 general ophthalmology questions and 400 paired answers were created by our ophthalmologists to represent commonly asked questions in real-world, across the spectrum of cataracts, myopia, and retinal diseases. This dataset was divided into finetuning (368 QnA pairs; 92%), and testing (40 QnA pairs; 8%). We find-tuned 5 different LLMs, including LLAMA2-7b, LLAMA2-7b-Chat, LLAMA2-13b, and LLAMA2-13b-Chat, based on the fine-tuning dataset as domain-specific knowledge. For the independent testing dataset, an additional 8 glaucoma QnA pairs were included. 200 responses to the testing dataset were generated by 5 fine-tuned LLMs for evaluation. A customized clinical evaluation rubric was used to guide GPT-4 evaluation of these LLM-generated responses, grounded on clinical accuracy, relevance, patient safety, and ease of understanding. GPT-4 evaluation was then compared against human ranking by 5 clinicians for clinical alignment.

§ The Valyu brief

Reading the full paper and taking notes. This takes a few seconds…

§ Ask this paper

Ask a question about this paper

Valyu reads the full text and answers from what the paper actually says.

Q.

Searching the other archives…

Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4 · Paraphernalia