An interdisciplinary, randomized, single-blind expert evaluation of state-of-the-art large language models for their implications and risks in medical diagnosis and management
Chen, P.; Cai, J.-f.; Zhou, J.; Chen, S.; Xu, C.; Yuan, L.; Dai, X.; Chen, X.; Wei, Y.; Li, X.; Gong, S.; Liang, X.; Yang, J.; Jin, J.; Dai, K.; Cui, Y.; Kuang, G.-M.; Xie, J.; Luo, L.; Xiao, H.; Yin, S.; Yang, J.; Yan, Y.; Chen, J.; Chen, Y.; Zhang, Q.; Zhou, Q.; Zhao, L.; Wu, M.; Tang, X.; Rong, L.; Wang, Z.; Qiu, W.; Wang, Y.; Cui, L.; Li, X.; Hu, Y.; Tao, H.; Wu, N.; Pai, P.; Wei, M.; To, M. K.-t.; Cheung, K. M. C.
Show abstract
BackgroundState-of-the-art (SOTA) large language models (LLMs) are poised to revolutionize clinical medicine by transforming diagnostic, therapeutic, and interdisciplinary reasoning. Despite their promising capabilities, rigorous benchmarking of these models is essential to address concerns about their clinical proficiency and safety, particularly in high-risk environments. MethodsThis study implemented a multi-disciplinary, randomized, single-blind evaluation framework involving 27 experienced specialty clinicians with an average of 25.9 years of practice. The assessment covered 685 simulated and real clinical cases across 13 subspecialties, including both common and rare conditions. Evaluators rated LLM responses on medical strength (0-10 scale, where >9.5 signified leading expert proficiency) and hallucination severity (0-5 scale for fabricated or misleading medical elements). Seven SOTA LLMs were tested, including top-ranked models from the ARENA leaderboard, with statistical analyses applied to adjust for confounders such as response length. FindingsThe evaluation revealed clinical plausibility in general-purpose LLMs, with Gemini 2.0 Flash leading raw scores and DeepSeek R1 excelling in adjusted analyses. Top models demonstrated proficiency comparable to a physician of 6 years post qualification experience (score [~]6.0), yet significant risks were noted. Instances of incompetence (scores [≤]4) were detected across specialties, and 40 hallucination instances involving fabricated conditions, medications, and classification errors. These findings underscore the importance of implementing stringent safeguards to mitigate potential adverse outcomes in clinical applications. InterpretationWhile SOTA LLMs show substantial promise in enhancing clinical reasoning and decision-making, their unguarded application in medicine could present serious risks, such as misinformation and diagnostic errors. Human expert oversight remains crucial, particularly given reported incompetence and hallucination risks. Larger, multi-center studies are warranted to evaluate their real-world performance and track their evolution before broader clinical adoption.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 94%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 92%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 92%
Similar papers in this journal
Similar papers in this journal
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 94%
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 92%
- Characterizing Long COVID: Deep Phenotype of a Complex Condition 92%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 90%
- ChatGPT in glioma patient adjuvant therapy decision making: ready to assume the role of a doctor in the tumour board? 89%
- Evaluating algorithmic fairness in the presence of clinical guidelines: the case of atherosclerotic cardiovascular disease risk estimation 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.