Why Large Language Models' Clinical Reasoning Fails: Insights from Explainable Deep Learning
Modi, M.; Krull, J. E.; Johnson, D.; Wang, X.; Gauntner, T. D.; Li, M.; Cheng, H.; Ma, A.; Zhang, P.; Stover, D. G.; Li, Z.; Ma, Q.
Show abstract
Medical large language models (LLMs) achieving high benchmark accuracy exhibit unexplained variability in clinical tasks, producing errors that clinicians cannot safeguard against. We evaluated clinical reasoning stability in GPT-5, MedGemma-27B-Text-IT, and OpenBioLLM-Llama3-70B using 355 systematic perturbations of physician-validated oncology cases and trained sparse autoencoders on 1 billion tokens from 50,000 MIMIC-IV clinical notes to decompose their internal representation. We find models exhibit dramatic reasoning instability, shifting staging accuracy by over 50% based solely on prompt format, or generating definitive staging in clinically insufficient scenarios. Sparse autoencoder analysis revealed hierarchical encoding in MedGemma, where high-magnitude features encode lexical identity and low-magnitude features encode contextual meaning. OpenBioLLM distributes information uniformly. We demonstrate these internal encoding structures differentially affect retrieval interventions, suggesting interventions effective for one architecture may harm another. We recommend healthcare institutions implement architecture-specific safety validation, as benchmark equivalence does not imply functional equivalence, with implications for AI safety beyond healthcare.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 95%
- Zero Shot Health Trajectory Prediction Using Transformer 95%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 95%
Similar papers in this journal
Similar papers in this journal
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 95%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 93%
- Weakly-Supervised Tumor Purity Prediction FromFrozen H&E Stained Slides 92%
Similar papers in this journal
Similar papers in this journal
- Application of Generative Artificial Intelligence to Utilise Unstructured Clinical Data for Acceleration of Inflammatory Bowel Disease Research 93%
- Multimodal surveillance of SARS-CoV-2 at a university enables development of a robust outbreak response framework 89%
- Internal checkpoint regulates T cell neoantigen reactivity and susceptibility to PD1 blockade 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.