Evaluating Clinical Foundation Models for Early Alzheimer's Disease and Related Dementia Prediction from Longitudinal EHRs
Farzana, S.; Arian, A.; Rundek, T.; Desvarieux, M.; Ahsan, H.
Show abstract
Early identification of Alzheimer's disease and related dementias (ADRD) remains challenging despite its importance for timely intervention, management of modifiable risk factors, and care planning. We developed and evaluated ADRD onset prediction models using longitudinal electronic health records (EHRs) from the All of Us Research Program at clinically meaningful lead times of 6, 12, 24, and 36 months before diagnosis, benchmarking interpretable count-based representations against four publicly available pretrained clinical foundation models (CLMBR-T, GPT-style, LLaMA-style, and Mamba) across multiple ADRD phenotype definitions. Count-based models consistently achieved the highest discrimination and calibration across all cohorts and prediction horizons. Predictive performance declined with increasing lead time for all approaches; however, the performance gap between count-based and pretrained representations progressively narrowed, with foundation models achieving comparable AUROC of 0.719 (compared to the AUROC of 0.738 of count-based model) at the 36-month horizon while providing higher sensitivity and F1 scores under a fixed operating threshold. External validation with zero-shot evaluation on UChicago EHRs exhibited limited generalizability for count-based and pretrained clinical foundation model based representations. These findings demonstrate that transparent count-based EHR representations remain the strongest overall approach for ADRD onset prediction, while pretrained clinical foundation models provide complementary advantages for long-term risk identification and establish a benchmark for evaluating transferable clinical representations in temporal ADRD risk prediction.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using Machine Learning and Electronic Health Record (EHR) Data for the Early Prediction of Alzheimer’s Disease and Related Dementias 94%
- Continuous Associations Between Remote Self-Administered Cognitive Measures and Imaging Biomarkers of Alzheimer’s Disease 93%
- Subjective cognition trajectories, Alzheimer biomarkers, and incident mild cognitive impairment 92%
Similar papers in this journal
- Machine Learning Prediction of Incidence of Alzheimers Disease Using Large-Scale Administrative Health Data 95%
- Interpretable deep learning approach for extracting cognitive features from hand-drawn images of intersecting pentagons in older adults 94%
- A scoping review of remote and unsupervised digital cognitive assessments in preclinical Alzheimer’s disease 93%
Similar papers in this journal
- Stratification of Alzheimer's Disease Patients Using Knowledge-Guided Unsupervised Latent Factor Clustering with Electronic Health Record Data 96%
- Spatial Navigation as a Digital Marker for Clinically Differentiating Cognitive Impairment Severity 96%
- Leveraging electronic health records to examine differential clinical outcomes in people with Alzheimer's Disease 95%
Similar papers in this journal
- Comparison and aggregation of event sequences across ten cohorts to describe the consensus biomarker evolution in Alzheimer’s disease 94%
- Development and validation of a harmonized memory score for multicenter Alzheimer's disease and related dementia research 94%
- Early Syndecan-4 Upregulation Predicts Cognitive and Pathological Trajectories in Alzheimer Disease 94%
Similar papers in this journal
- Quantitative longitudinal predictions of Alzheimer's disease by multi-modal predictive learning 94%
- Profiles of cognitive change in preclinical Alzheimer's disease using change-point analysis 93%
- Blood Biomarkers for Diagnosis & Differential Diagnosis of Alzheimers Disease in Real-World Clinical Populations: A Systematic Review 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.