Back

Do Language Models Think Like Doctors?

McCoy, L. G.; Swamy, R.; Sagar, N.; Wang, M.; Cao, J.; Bacchi, S.; Fong, N.; Tan, N. C. K.; Tan, K.; Buckley, T.; Brodeur, P.; Celi, L. A.; Manrai, A. K.; Humbert, A. J.; Rodman, A.

2025-02-12 health informatics
10.1101/2025.02.11.25321822 medRxiv
Show abstract

BackgroundWhile large language models (LLMs) are being increasingly deployed for clinical decision support, existing evaluation methods like medical licensing exams fail to capture critical aspects of clinical reasoning including reasoning in dynamic clinical circumstances. Script Concordance Testing (SCT), a decades-old medical assessment tool, offers a nuanced way to assess how new information influences diagnostic and therapeutic decisions under uncertainty. MethodsWe developed a comprehensive and publicly available benchmark comprising 750 SCT questions from 10 internationally diverse medical datasets--9 previously unreleased--spanning multiple specialties and institutions. Each question presents a clinical scenario then asks how new information affects the likelihood of a diagnosis or management decision, scored against expert panels (Figure 1). We evaluated four state-of-the-art LLMs against the combined responses of 1070 medical students, 193 resident physicians, and 300 attending physicians in total across all datasets. O_FIG O_LINKSMALLFIG WIDTH=191 HEIGHT=200 SRC="FIGDIR/small/25321822v1_fig1.gif" ALT="Figure 1"> View larger version (46K): org.highwire.dtl.DTLVardef@b0ebc4org.highwire.dtl.DTLVardef@146aef6org.highwire.dtl.DTLVardef@188b3cdorg.highwire.dtl.DTLVardef@1d47229_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFIGURE 1.C_FLOATNO Example of a Script Concordance Test with Scoring Schema. 1A demonstrates formatting for students, 1B demonstrates formatting for LLM use, as well as an example of the scoring structure. C_FIG ResultsLLMs demonstrated markedly lower performance on SCTs compared to their typical achievement on medical multiple choice benchmarks. GPT-4o achieved the highest performance (63.6% {+/-} 1.2%), significantly outperforming other models (Claude-3.5 Sonnet: 58.8% {+/-} 1.2%, o1-preview: 58.5% {+/-} 1.3%, Gemini-1.5-Pro: 54.4% {+/-} 1.4%). Models matched or exceeded student performance on multiple examinations, but did not reach the level of senior residents or attending physicians (Figure 2). Surprisingly, the integrated-chain-of-thought o1-preview model underperformed compared to GPT-4o, a contrast with their relative performance on other medical benchmarks. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=122 SRC="FIGDIR/small/25321822v1_fig2.gif" ALT="Figure 2"> View larger version (10K): org.highwire.dtl.DTLVardef@930cbeorg.highwire.dtl.DTLVardef@299ea7org.highwire.dtl.DTLVardef@6ef14forg.highwire.dtl.DTLVardef@1a47cd5_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 2.C_FLOATNO Overall Performance of Models on Full SCT Benchmark. Scores reported as mean percentage of maximum possible score, +/-standard error of mean (SEM). Statistical analysis performed with repeated measures Friedman test with post-hoc Conover test with Bonferroni correction applied. *p<0.05, **p<0.01, ***p<0.001 C_FIG ConclusionsSCT represents a challenging and distinctive benchmark for evaluating LLM clinical reasoning capabilities, revealing limitations not apparent in traditional MCQ-based assessments. This work demonstrates the value of SCT in providing a more nuanced evaluation of medical AI systems and highlights specific areas where current models may fall short in clinical reasoning tasks. We are making our benchmark publicly available in a secure format to foster collaborative improvement of clinical reasoning capabilities in LLMs. Brief SummaryWhile large language models excel at traditional medical knowledge tests, their performance on our new public Script Concordance Test benchmark reveals important limitations in clinical reasoning capabilities, particularly in processing new information under uncertainty.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.