Back

Medical errors in large language models revealed using 1,000 synthetic clinical transcripts

Auger, S. D.; Scott, G.

2026-03-25 health informatics
10.64898/2026.03.23.26349082 medRxiv
Show abstract

Current clinical evaluations of large language models (LLMs) rely on datasets which fail to reflect real-world medical complexity. We developed a high-throughput simulation of patients presenting with headache and generated 1,000 doctor-patient transcripts, enabling an unprecedented mapping of clinical reasoning failures across a vast spectrum of demographic and clinical phenotypes. While GPT-5.2 achieved 97.5% diagnostic accuracy with full histories (95% confidence interval: 95.0-99.5), incomplete transcripts triggered hazardous recommendations. Instead of requesting further information, both models discouraged essential investigations, including 100% of lumbar punctures in subarachnoid haemorrhage cases. For life- or sight-threatening emergencies, models inappropriately downgraded triage to self-management or routine follow-up in up to 54.8% cases (GPT-5-mini; 95% CIs: 40.5-69.0). GPT-5.2 triage was significantly less safe for females than males (odds ratio 3.2 95% CI 1.4-7.1). Our methodology transforms medical AI evaluation from simple snapshots into a comprehensive safety stress-test, revealing critical failures in risk calibration despite high diagnostic accuracy.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.