Evaluating the Generalizability of EEG-Based AI Models in Alzheimers and Dementia Diagnosis
Saini, R.; Simistira Liwicki, F.; Rakesh, S.; Mokayed, H.; Acharya, S.; Singh, D.; Gupta, V.; Arpak, E. S.; Chakladar, D. D.
Show abstract
INTRODUCTIONWe thoroughly investigated the generalizability of deep learning models trained on electroencephalography (EEG) data to detect Alzheimers disease and dementia at the individual subject level. Although average model performance appears strong, it may obscure large inter-individual variability, raising concerns for clinical deployment. METHODSWe trained a Hopfield-enhanced deep neural network on a publicly available EEG dataset consisting of 88 participants, including individuals diagnosed with Alzheimers disease (AD), frontotemporal dementia (FTD), and cognitively normal controls (CN). Resting-state EEG recordings were segmented and used to train the model in a leave-onesubject-out (LOSO) cross-validation setup across multiple detection tasks: AD vs. CN, AD vs. FTD, FTD vs. CN, and AD vs. FTD vs. CN. RESULTWhile the model demonstrated high average performance (e.g., up to 83% accuracy), subject-level results revealed inconsistencies. Some individuals achieved perfect prediction even at the first training epoch, suggesting spurious memorization, while others predicted falsely throughout, with performance below chance. These patterns persisted despite consistent training conditions and no data leakage. DISCUSSIONOur findings highlight that strong group-level performance may be misleading in clinical settings, where decisions are made at the individual level. The models should be generalizable across individuals and be evaluated per individual before being considered for diagnostic use. Hopfield networks show promise in capturing patterns in EEG data, but patient-level validation and transparent reporting are essential to avoid premature clinical translation. HighlightsO_LIDeep learning models trained on EEG can achieve high average performance in detecting Alzheimers and dementia disease. C_LIO_LIThe subject-level evaluation revealed significant variability, including below-chance performance for some indi- viduals despite overall strong results. C_LIO_LIGroup-level metrics alone may be misleading; General- izable models and individual-level validation are critical before the clinical adoption of EEG-based AI models. C_LI
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable deep learning approach for extracting cognitive features from hand-drawn images of intersecting pentagons in older adults 93%
- A Machine-Learning Based Objective Measure for ALS Disease Severity 92%
- Quantifying Device Type and Handedness Biases in a Remote Parkinson’s Disease AI-Powered Assessment 92%
Similar papers in this journal
- c-Triadem: A constrained, explainable deep learning model to identify novel biomarkers in Alzheimer’s disease 94%
- Random forest model for feature-based Alzheimer's disease conversion prediction from early mild cognitive impairment subjects 94%
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 93%
Similar papers in this journal
- Uncertainty in Deep Learning for EEG under Dataset Shifts 96%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 92%
- Subtle anomaly detection in MRI brain scans: Application to biomarkers extraction in patients with de novo Parkinson’s disease 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.