Automating Evaluation of LLM-generated Responses to Patient Questions about Rare Diseases
Zhao, M.; Oh, I. Y.; Gupta, A.; Cohen-Cutler, S.; Harmoney, K. M.; Lai, A. M.; Sisk, B. A.
Show abstract
ObjectivesPatients with rare diseases often struggle to find accurate medical information, and large language model (LLM)-based chatbots may help meet this need. However, evaluating LLM-generated free-text answers typically requires physician review, which is time-consuming and difficult to scale. This study compared traditional natural language processing (NLP) metrics to emerging LLM-based evaluation approaches for assessing answer quality in the context of Complex Lymphatic Anomalies (CLAs). Materials and MethodsWe compiled 25 common patients questions about CLAs and generated 175 responses to these questions from seven LLMs. Three expert physicians scored these responses for accuracy. We compared these physician-assigned scores with automated scores, generated by four NLP sentence similarity metrics (BLEU, ROUGE, METEOR, BERTScore) and six LLM evaluators (GPT-4, GPT-4o, Qwen3-32B, DeepSeek-R1-14B, Gemma3-27B, LLaMA3.3-70B). We examined both LLM-based scoring with and without reference answers (reference-guided vs. reference-free). We calculated Spearman, Phi, and Kendalls Tau correlation coefficients to assess alignment between automated and physician-assigned scores. ResultsLLM-based evaluation demonstrated stronger alignment with physician-assigned scores than NLP metrics. The reference-guided GPT-4 evaluator achieved the highest correlation with physician-assigned scores ({rho}=0.758), followed by GPT-4o ({rho}=0.727). NLP metrics showed weak to moderate correlations with physician-assigned scores ({rho}=0.240-0.403). Reference-guided scoring outperformed reference-free methods. DiscussionReference-guided LLM-based evaluation methods approximate expert physicians judgment better than traditional NLP metrics, offering an effective, scalable approach for assessing LLM-generated responses to patient questions about rare disease. ConclusionLLM-based evaluation, particularly reference-guided scoring with GPT models, can support the scalable development and evaluation of LLM-based rare disease-specific chatbot systems.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 95%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 95%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 94%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.