Comparative Benchmarking of Five Contemporary Language Models on Clinical Reasoning
Al-Risheq, A. N.
Show abstract
BackgroundThe rapid integration of Large Language Models (LLMs) into healthcare raises critical questions regarding their safety and reliability. While models often score highly on standardized medical examinations, their performance in open-ended, high-stakes clinical decision-making, particularly when navigating strict safety contraindications, remains under-explored (Xiao et al., 2025; Wu et al., 2024). ObjectiveThis study benchmarks five contemporary "reasoning" models ChatGPT-5.2 (Thinking), Kimi K2 Thinking, DeepSeek V3.2 deepthink, Gemini 3 Pro, and Claude 4.5 Opus (Thinking) on diagnostic accuracy, management appropriateness, and adherence to safety protocols. MethodsI designed 15 synthetic clinical vignettes covering diverse medical specialties, including a targeted "safety trap" scenario involving severe penicillin anaphylaxis. I manually evaluated model responses against a gold-standard answer key using a strict scoring rubric that penalized unsafe recommendations regardless of diagnostic accuracy. ResultsKimi K2 Thinking and ChatGPT-5.2 achieved the highest aggregate scores (3.50/3.50), demonstrating 100% diagnostic accuracy and perfect safety adherence. DeepSeek V3.2 followed closely (3.46). Conversely, Gemini 3 Pro and Claude 4.5 Opus incurred significant safety penalties for suggesting carbapenems in a patient with severe IgE-mediated anaphylaxis, a violation of the studys strict safety rubric, despite otherwise high clinical competence. ConclusionMy analysis reveals that while modern reasoning (Chain-of-Thought) models possess exceptional diagnostic capabilities, they differ significantly in their handling of "hard" safety constraints. Models that prioritize conservative heuristics (Kimi, GPT-5.2) outperformed those that attempted more nuanced but risky pharmacological justifications (Gemini, Opus) in this specific safety benchmarking context (Large Language Models Lack Essential Metacognition for Reliable Medical Reasoning, 2024).
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 91%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 91%
Similar papers in this journal
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 93%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 91%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 91%
Similar papers in this journal
- Refining LLMs Outputs with Iterative Consensus Ensemble (ICE) 93%
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 90%
- Predicting the physiological effects of multiple drugs using electronic health record 90%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 93%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 92%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.