Automated Interpretation of EEG Reports Using a Large Language Model with Structured Confidence Outputs
Tian, W.; Bergner, S.; Moiseev, A.; Popowich, F.; Medvedev, G.; Richardson, M. P.; Rodionov, R.; Xi, P.; Doesburg, S. M.; Ribary, U.; Winston, J. S.; Vakorin, V. A.
Show abstract
Background: Free-text EEG reports typically lack structure, hindering scalable analysis. We evaluate a large language model (LLM) pipeline to extract structured diagnostic labels and confidence levels from these reports. Methods: We developed a hierarchical annotation schema to classify EEG reports for four specific abnormality types using a four-point confidence scale. To establish ground truth, two certified EEG technicians annotated a diverse dataset of reports authored by neurologists with distinct writing styles. We then implemented a grammar-constrained Mistral-7B pipeline, iteratively prompt-tuned on a development set to mirror these expert annotations. The pipeline's effectiveness was evaluated against the human expert benchmark using core agreement (diagnostic accuracy) and certainty-adjusted agreement (confidence alignment), with classical NLP models serving as a secondary baseline. Results: Mistral-7B significantly outperformed baselines, achieving 96% accuracy for overall abnormality detection, approaching the human benchmark of 98%. Crucially, the model successfully identified rare epileptiform abnormalities where traditional models failed and generalized robustly across distinct reporting styles. While diagnostic accuracy was high, a performance gap persisted in certainty-adjusted agreement, indicating that accurately modeling nuanced clinical confidence remains a challenge. Conclusion: LLMs can effectively automate the extraction of structured diagnostic information from EEG reports with near-human accuracy and strong generalization. While confidence calibration requires further refinement, the combination of accurate classification and explainability makes this pipeline a promising tool for standardizing clinical data at scale. Keywords: Routine Clinical Electroencephalography; Large Language Models; Clinical NLP; Confidence Assessment; Explainable AI; Neurophysiological Evaluation
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 91%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 90%
- Automating literature screening and curation with applications to computational neuroscience 90%
Similar papers in this journal
- Uncertainty in Deep Learning for EEG under Dataset Shifts 94%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 92%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 90%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 91%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 91%
Similar papers in this journal
- Evaluating the generalisability of region-naïve machine learning algorithms for the identification of epilepsy in low-resource settings 91%
- Longitudinally Tracking Personal Physiomes for Precision Management of Childhood Epilepsy 90%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 90%
Similar papers in this journal
- NLP-based tools for localization of the Epileptogenic Zone in patients with drug-resistant focal epilepsy 93%
- Online speech synthesis using a chronically implanted brain-computer interface in an individual with ALS 92%
- Recurrent Neural Network-based Acute Concussion Classifier using Raw Resting State EEG Data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.