Back

Large Language Model - Enhanced Decision Tree Framework for Identifying Multiple Sclerosis Diagnoses from Clinical Documentation

Venkatesh, S.; DelSignore, M.; Wu, X.; Morris, M.; Kerr, W. T.; Visweswaran, S.; Wang, Y.; Xia, Z.

2026-07-17 neurology
10.64898/2026.07.14.26357416 medRxiv
Show abstract

Background. Early diagnosis and intervention are crucial in multiple sclerosis (MS), yet diagnostic delays are common. Large language models (LLMs) such as generative pre-trained transformers (GPTs) may help streamline diagnostic workflows by extracting MS diagnostic signals from clinical notes. Objective. To derive MS diagnosis status from the first neurology note using a computable algorithm based on the 2017 McDonald criteria and applying GPT-4 for node-level reasoning within a structured decision framework. Methods. We analyzed first neurology notes from 125 randomly selected patients (including those with MS, related disorders, and controls) enrolled in a clinic cohort between 2017 and 2023. We included the clinical history and diagnostic testing sections but redacted the assessment and plan. We converted the 2017 McDonald criteria into a decision tree and provided expert-curated clinical knowledge to guide GPT-4 reasoning at each decision node. GPT-4 generated binary decisions at each node to traverse the tree and classified MS diagnoses at terminal nodes. We evaluated performance against neurologist-assessed diagnoses and characterized hallucinations (non-factual, incongruent, irrelevant, over-reliant, and logical reasoning errors). Results. In this study cohort (mean age 40{+/-}13 years; 81% women) representative of the clinic population, GPT-4 performed well in predicting MS diagnosis (84% accuracy, 79% precision, 74% recall, 91% specificity) using first neurology notes. Hallucinations occurred in 32 cases (26%), most commonly incoherence (75%) and overreliance (47%). Conclusion. A structured, LLM-guided decision framework can flag MS diagnoses from early clinical documentation. Large-scale studies are needed to mitigate hallucinations, validate this approach, and test implementation in clinical settings.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.