Back

Large Language Model - Enhanced Decision Tree Framework for Identifying Multiple Sclerosis Diagnoses from Clinical Documentation

Venkatesh, S.; DelSignore, M.; Wu, X.; Morris, M.; Kerr, W. T.; Visweswaran, S.; Wang, Y.; Xia, Z.

2026-07-17 neurology
10.64898/2026.07.14.26357416 medRxiv
Show abstract

Background. Early diagnosis and intervention are crucial in multiple sclerosis (MS), yet diagnostic delays are common. Large language models (LLMs) such as generative pre-trained transformers (GPTs) may help streamline diagnostic workflows by extracting MS diagnostic signals from clinical notes. Objective. To derive MS diagnosis status from the first neurology note using a computable algorithm based on the 2017 McDonald criteria and applying GPT-4 for node-level reasoning within a structured decision framework. Methods. We analyzed first neurology notes from 125 randomly selected patients (including those with MS, related disorders, and controls) enrolled in a clinic cohort between 2017 and 2023. We included the clinical history and diagnostic testing sections but redacted the assessment and plan. We converted the 2017 McDonald criteria into a decision tree and provided expert-curated clinical knowledge to guide GPT-4 reasoning at each decision node. GPT-4 generated binary decisions at each node to traverse the tree and classified MS diagnoses at terminal nodes. We evaluated performance against neurologist-assessed diagnoses and characterized hallucinations (non-factual, incongruent, irrelevant, over-reliant, and logical reasoning errors). Results. In this study cohort (mean age 40{+/-}13 years; 81% women) representative of the clinic population, GPT-4 performed well in predicting MS diagnosis (84% accuracy, 79% precision, 74% recall, 91% specificity) using first neurology notes. Hallucinations occurred in 32 cases (26%), most commonly incoherence (75%) and overreliance (47%). Conclusion. A structured, LLM-guided decision framework can flag MS diagnoses from early clinical documentation. Large-scale studies are needed to mitigate hallucinations, validate this approach, and test implementation in clinical settings.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.3%
19.1%
2
Frontiers in Neurology
102 papers in training set
Top 0.1%
15.6%
3
Multiple Sclerosis Journal
21 papers in training set
Top 0.1%
8.1%
4
Multiple Sclerosis and Related Disorders
15 papers in training set
Top 0.1%
6.5%
5
Journal of Neurology, Neurosurgery & Psychiatry
30 papers in training set
Top 0.1%
5.3%
50% of probability mass above
6
Annals of Clinical and Translational Neurology
34 papers in training set
Top 0.2%
4.2%
7
BMC Neurology
14 papers in training set
Top 0.1%
3.5%
8
Frontiers in Digital Health
24 papers in training set
Top 0.4%
3.3%
9
Communications Medicine
113 papers in training set
Top 1%
2.9%
10
PLOS ONE
5266 papers in training set
Top 42%
2.5%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
2.0%
12
PLOS Digital Health
106 papers in training set
Top 2%
2.0%
13
eBioMedicine
183 papers in training set
Top 4%
1.2%
14
Scientific Reports
3612 papers in training set
Top 68%
1.1%
15
Journal of Neurology
28 papers in training set
Top 0.8%
1.1%
16
Journal of Translational Medicine
57 papers in training set
Top 2%
1.0%
17
Neurology Neuroimmunology & Neuroinflammation
12 papers in training set
Top 0.1%
1.0%
18
Journal of the Neurological Sciences
18 papers in training set
Top 0.5%
0.9%
19
Neurology
50 papers in training set
Top 1%
0.9%
20
Orphanet Journal of Rare Diseases
21 papers in training set
Top 0.5%
0.9%
21
Annals of Neurology
64 papers in training set
Top 1%
0.9%
22
Brain Communications
166 papers in training set
Top 3%
0.9%
23
Clinical Immunology
21 papers in training set
Top 0.3%
0.9%
24
BMJ Open
601 papers in training set
Top 13%
0.6%
25
European Journal of Neurology
22 papers in training set
Top 0.8%
0.6%
26
BMC Medicine
176 papers in training set
Top 5%
0.6%