Medical errors in large language models revealed using 1,000 synthetic clinical transcripts
Auger, S. D.; Scott, G.
Show abstract
Current clinical evaluations of large language models (LLMs) rely on datasets which fail to reflect real-world medical complexity. We developed a high-throughput simulation of patients presenting with headache and generated 1,000 doctor-patient transcripts, enabling an unprecedented mapping of clinical reasoning failures across a vast spectrum of demographic and clinical phenotypes. While GPT-5.2 achieved 97.5% diagnostic accuracy with full histories (95% confidence interval: 95.0-99.5), incomplete transcripts triggered hazardous recommendations. Instead of requesting further information, both models discouraged essential investigations, including 100% of lumbar punctures in subarachnoid haemorrhage cases. For life- or sight-threatening emergencies, models inappropriately downgraded triage to self-management or routine follow-up in up to 54.8% cases (GPT-5-mini; 95% CIs: 40.5-69.0). GPT-5.2 triage was significantly less safe for females than males (odds ratio 3.2 95% CI 1.4-7.1). Our methodology transforms medical AI evaluation from simple snapshots into a comprehensive safety stress-test, revealing critical failures in risk calibration despite high diagnostic accuracy.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 94%
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 94%
- Conformal prediction enables disease course prediction and allows individualized diagnostic uncertainty in multiple sclerosis 93%
Similar papers in this journal
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 95%
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 92%
- Attributes and predictors of Long-COVID: analysis of COVID cases and their symptoms collected by the Covid Symptoms Study App 90%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 93%
- Automated identification of abnormal infant movements from smart phone videos 92%
- Community-acquired pneumonia identification from electronic health records in the absence of a gold standard: a Bayesian latent class analysis 91%
Similar papers in this journal
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 92%
- Clinical trial emulation can identify new opportunities to enhance the regulation of drug safety in pregnancy 91%
- Adherence and sustainability of interventions informing optimal control against COVID-19 pandemic 90%
Similar papers in this journal
- Benchmarking transformer-based models for medical record deidentification: A single centre, multi-specialty evaluation 94%
- A graph-embedded topic model enables characterization of diverse pain phenotypes among UK Biobank individuals 94%
- EYE-Llama, an in-domain large language model for ophthalmology 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.