Assessing AI and Neurologist Diagnostic Reasoning Against Neuropathological Ground Truth
Leng, Y.; Noori, A.; Dickson, J. R.; Serrano-Pozo, A.; Avetisyan, M.; Rodriguez, D.; Rosenberg, E. S.; He, Y.; Oakley, D. H.; Khurana, V. S.; Hyman, B. T.; Frosch, M. P.; Das, S.
Show abstract
BACKGROUND Accurate differential diagnosis of complex neurological disorders remains challenging due to overlapping clinical features and heterogeneous disease presentations. Although large language models (LLMs) show promise in clinical reasoning, prior studies benchmark performance against clinician consensus rather than biological ground truth. A neuropathologically confirmed benchmark dataset for evaluating diagnostic AI in neurology is currently lacking. METHODS We introduce NeuroBench, a curated benchmark of complex neurological cases with neuropathologically confirmed gold-standard diagnoses, and DIAGNO, a confidence-aware LLM-based system for neurological diagnosis. NeuroBench comprises 203 retrospective case summaries from the Massachusetts General Hospital Brain Cutting Conference with corresponding autopsy-confirmed diagnoses. DIAGNO generated top-3 differential diagnoses, employing retrieval-augmented generation (RAG) for lower-confidence cases. Performance was assessed by three independent blinded adjudicators who evaluated both DIAGNO and neurologists against neuropathological ground truth. RESULTS NeuroBench encompassed 79 unique neuropathological diagnoses, spanning conditions including cerebrovascular disease, brain tumors, neurological infections, and various neurodegenerative and inflammatory disorders. DIAGNO matched or outperformed neurologists in top-3 accuracy (0.67 versus 0.63) and taxonomy-level accuracy (0.74 versus 0.66). In cases of disagreement, DIAGNO was more often correct than neurologists (29 versus 19 cases). Diagnostic concordance between DIAGNO and neurologists was high (90% agreement in top-3 predictions), even when both were incorrect, suggesting strong alignment in diagnostic reasoning. On NeuroBench, DIAGNO also outperformed GPT-4o baseline and DeepSeek R1 across all top-k accuracy metrics. In a real-world evaluation on eight complex cases with differentials from Mass General Brigham, neurologists rated DIAGNO's reasoning favorably (mean 4.03/5) across multiple dimensions of clinical utility and safety. CONCLUSIONS NeuroBench establishes neuropathological confirmation as the appropriate standard for evaluating diagnostic AI in neurology, moving beyond clinician-referenced benchmarking to define the ceiling of diagnostic accuracy. Evaluated against this standard, DIAGNO achieved expert-level diagnostic performance and received favorable clinician ratings in real-world applications, supporting its potential as a clinical decision-support tool in neurology.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Conformal prediction enables disease course prediction and allows individualized diagnostic uncertainty in multiple sclerosis 92%
- Interpretable deep learning approach for extracting cognitive features from hand-drawn images of intersecting pentagons in older adults 92%
- High-Sensitivity Pan-Cancer AI Assessment of Lymph Node Metastasis via Uncertainty Quantification 91%
Similar papers in this journal
Similar papers in this journal
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 92%
- Enhancing Early Detection of Cognitive Decline in the Elderly through Ensemble of NLP Techniques: A Comparative Study Utilizing Large Language Models in Clinical Notes 90%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 89%
Similar papers in this journal
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 90%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 90%
- LUNAR: A Deep Learning Model to Predict Glioma Recurrence Using Integrated Genomic and Clinical Data 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.