Automated Evaluation of Large Language Model Response Concordance with Human Specialist Responses on Physician-to-Physician eConsult Cases
Wu, D. J.; Haredasht, F. N.; Wu, D.; Ravi, V.; McCoy, L. G.; Weng, Y.; Chopra, K.; Everett, S.; Nageeb, G.; Chen, W.; Ma, S.; Maharaj, S. K.; Tran, J.; Rosengaus, L.; Giang, L.; Jee, O.; Goh, E.; Chen, J. H.
Show abstract
Specialist consults in primary care and inpatient settings typically address complex clinical questions beyond standard guidelines. eConsults have been developed as a way for specialist physicians to review cases asynchronously and provide clinical answers without a formal patient encounter. Meanwhile, large language models (LLMs) have approached human-level performance on structured clinical tasks, but their real-world effectiveness requires evaluation, which is bottlenecked by time-intensive manual physician review. To address this, we evaluate two automated methods: LLM-as-judge and a decompose-then-verify framework that breaks down AI answers into verifiable claims against human eConsult responses. Using 40 real-world physician-to-physician eConsults, we compared AI-generated responses to human answers using both physician raters and automated tools. LLM-as-judge outperformed decompose-then-verify, achieving human-level concordance assessment with F1-score of 0.89 (95% CI: 0.750, 0.960) and Cohens kappa of 0.75 (95% CI 0.47,0.90) --comparable to physician inter-rater agreement {kappa} = 0.69-0.90 (95% CI 0.43-1.0).
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
Similar papers in this journal
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.