Unmasking True Clinical Competence: The Importance of Adaptive and Open-Ended Evaluation for LLMs in Cardiology
Chao, C.-J.; Kumar, A.; Mishra, A.; Wang, Y.-C.; Tsai, C.-M.; Ali, N. B.; Sharma, S.; Farina, J. M.; Arsanjani, R.; Jiang, Y.; Lam, M.
Show abstract
BackgroundLarge language models (LLMs) achieve impressive accuracy on multiple-choice (MC) medical examinations, but this performance alone may not accurately reflect their clinical reasoning abilities. MC evaluations of LLMs risk inflating apparent competence by rewarding recall rather than genuine clinical adaptability, particularly in specialized medical domains such as cardiology. AimTo rigorously evaluate the clinical reasoning, contextual adaptability, and answer-reasoning consistency of a state-of-the-art reasoning-based LLM in cardiology using MC, open-ended, and clinically modified question formats. MethodsWe assessed GPT o1-preview using 185 board-style cardiology questions from the American College of Cardiology Self-Assessment Program (ACCSAP) in MC and open-ended formats. A subset of 66 questions underwent modifications of critical clinical parameters (e.g., ascending aorta diameter, ejection fraction) to evaluate model adaptability to context changes. The models answer and reasoning correctness were graded by cardiology experts. Statistical differences were analyzed using exact McNemar tests. ResultsGPT o1-preview demonstrated high baseline accuracy on MC questions (93.0% answers, 92.4% reasoning). Performance significantly decreased with open-ended questions (80.0% answers, 80.5% reasoning; p<0.001). For modified MC questions, accuracy decreased significantly (answers: 93.9% to 66.7%; reasoning: 93.9% to 71.2%; both p<0.001), as well as answer-reasoning concordance (93.9% to 66.7%, p<0.001). ConclusionsUsing existing MC question formats substantially overestimates the performance and clinical reasoning capabilities of GPT o1-preview. Incorporating open-ended, clinically adaptive questions and evaluating answer-reasoning concordance are essential for accurately assessing the real-world clinical decision-making competencies of LLMs in cardiology.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 94%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 93%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 92%
Similar papers in this journal
- Automated Diagnostic Reports from Images of Electrocardiograms at the Point-of-Care 93%
- Deep Learning-Based Multi-View Echocardiographic Framework for Comprehensive Diagnosis of Pericardial Disease 90%
- Automated Echocardiographic Detection of Mitral Valve Prolapse and Mitral Regurgitation with Video-based Artificial Intelligence Algorithms 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.