Pathology's Last Exam: Stress-Testing Diagnostic Reasoning and Safety in Large Language Models
Reitsam, N. G.; Gustav, M.; Jesinghaus, M.; Maerkl, B.; Foersch, S.; Kather, J. N.
Show abstract
Large language models (LLMs) are evolving into diagnostic co-pilots, yet current benchmarks fail to test the integrated, stepwise reasoning required in diagnostic pathology. Here, we present Pathologys Last Exam (PLE), a curated, highly detailed, text-based benchmark of 100 complex cases spanning organ systems, enriched for rare/challenging entities, plus 20 adversarial cases designed to stress-test model safety. Each case provides structured blocks (Primary, Clinical, Histopathology, IHC/Special Stains, Molecular Pathology) with stepwise information release mirroring real sign-out. We evaluated five LLMs (one proprietary, four open-source) across different stages. While the best model (GPT-5) achieved 70% accuracy on full evidence, performance on safety tests was alarming. Models frequently failed to detect biological contradictions, confidently diagnosing nonsensical "mix-up" cases rather than refusing them. This reveals a critical safety gap: high diagnostic capability is currently coupled with a dangerous inability to recognize impossible clinical scenarios. PLE provides a framework to measure and mitigate these risks before clinical deployment, as well as a foundation for developing multimodal evaluation protocols that can be extended to vision-language models and autonomous diagnostic agents in the future.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable multimodal deep learning for real-time pan-tissue pan-disease pathology search on social media 95%
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 94%
- Genomic Characterization of Lung Cancer in Never-Smokers Using Deep Learning 93%
Similar papers in this journal
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 95%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 93%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 92%
Similar papers in this journal
Similar papers in this journal
- Generalizing AI-driven Assessment of Immunohistochemistry across Immunostains and Cancer Types: A Universal Immunohistochemistry Analyzer 94%
- A Deep Learning Model for Molecular Label Transfer that Enables Cancer Cell Identification from Histopathology Images 94%
- Real-World Benchmarking and Validation of Foundation Model Transformers for Endometrial Cancer Subtyping from Histopathology 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.