Evaluating Large Language Model Diagnostic Performance on JAMA Clinical Challenges via a Multi-Agent Conversational Framework
Sangwon, K. L.; Zhang, J.; Steele, R.; Stryker, J.; Lee, J. V.; Choi, J.; Vishwanath, K.; Alber, D. A.; Kondziolka, D.; Mankowski, M.; Oermann, E. K.
Show abstract
Background & ObjectiveStandard clinical LLM benchmarks use multiple-choice vignettes that present all information up front, unlike real encounters where clinicians iteratively elicit histories and objective data. We hypothesized that such formats inflate LLM performance and mask weaknesses in diagnostic reasoning. We developed and evaluated a multi-AI agent conversational framework that converts JAMA Clinical Challenge cases into multi-turn dialogues, and assessed its impact on diagnostic accuracy across frontier LLMs. MethodsWe adapted 815 diagnostic cases from 1,519 JAMA Clinical Challenges into two formats: (1) original vignette and (2) multi-agent conversation with a Patient AI (subjective history) and a System AI (objective data: exam, labs, imaging). A Clinical LLM queried these agents and produced a final diagnosis. Models tested were O1 (OpenAI), GPT-4o (OpenAI), LLaMA-3-70B (Meta), and Deepseek-R1-distill-LLaMA3-70B (Deepseek), each in multiple-choice and free-response modes. Free-response grading used a separate GPT-4o judge for diagnostic equivalence. Accuracy (Wilson 95% CIs) and conversation lengths were compared using two-tailed tests. ResultsAccuracy decreased for all models when moving from vignettes to conversations and from multiple-choice to free-response (p<0.0001 for all pairwise comparisons). In vignette multiple-choice, accuracy was O1 79.8% (95% CI, 76.9%-82.4%), GPT-4o 74.5% (71.4%-77.4%), LLaMA-3 70.9% (69.5%-72.2%), Deepseek-R1 69.0% (67.5%-70.4%). In conversation multiple-choice: O1 69.1% (65.8%-72.2%), GPT-4o 51.3% (49.8%-52.8%), LLaMA-3 49.7% (48.2%-51.3%), Deepseek-R1 34.0% (32.6%-35.5%). In conversation free-response: O1 31.7% (28.6%-34.9%), GPT-4o 20.7% (19.5%-22.0%), LLaMA-3 22.9% (21.6%-24.2%), Deepseek-R1 9.3% (8.4%-10.2%). O1 generally required fewer conversational turns than GPT-4o, suggesting more efficient multi-turn reasoning. ConclusionsConverting vignettes into multi-agent, multi-turn dialogues reveals substantial performance drops across leading LLMs, indicating that static multiple-choice benchmarks overestimate clinical reasoning competence. Our open-source framework offers a more rigorous and discriminative evaluation and a realistic substrate for educational use, enabling assessment of iterative information-gathering and synthesis that better reflects clinical practice.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 95%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 90%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
Similar papers in this journal
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 91%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 91%
- Optimal policy determination in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: A patient-physician digital twin dyad with deep Q-learning for treatment selection 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.