Longitudinal Clinical Foundation Models Augmented with Genomics for Early Detection and Risk Stratification of Inherited Cardiomyopathy
Zolensky, A. L.; Kripke, C. M.; Keat, K.; Damrauer, S. M.; Levin, M. G.; Verma, A.
Show abstract
Hypertrophic and dilated cardiomyopathy (HCM and DCM) carry substantial morbidity and mortality, yet diagnosis may be delayed, particularly when presentation is nonspecific. Existing machine-learning approaches to cardiomyopathy phenotyping, genotype prediction, and risk stratification commonly rely on disease-specific, hand-engineered features drawn from echocardiography, cardiac MRI, ECG, or curated clinical variables. We evaluated whether a general-purpose clinical foundation model, CLMBR-T-base, pre-trained via next-clinical-event prediction with no cardiomyopathy-specific supervision, could produce linearly separable embeddings for all three case/control cohorts. Using EHR data from the Penn Medicine BioBank, we constructed cohorts for (1) prediction of a first recorded qualifying HCM/DCM diagnosis at 1-, 3-, and 6-month horizons, decomposed into eventual-versus-never-case and imminent-versus-eventual comparisons; (2) genetic carrier status prediction among diagnosed patients with completed gene panels; and (3) prediction of heart-failure hospitalization, and all-cause mortality as both binary and time-to-event outcomes. Linear probes fitted to frozen embeddings achieved AUROCs of 0.75-0.82 for onset prediction, 0.74-0.75 for genotype status, and Harrell's concordance of 0.65-0.80 for time-to-event outcomes. Decomposing the onset prediction task reveals that the model often misclassifies patients who were diagnosed later as positive, suggesting the patient journey embeddings encode disease state more reliably than care timing. These results suggest that a single, generically pretrained EHR embedding can support multiple clinically motivated prediction problems in CM without disease-specific feature engineering.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Exome-by-phenome-wide rare variant gene burden association with electronic health record phenotypes 91%
- Polygenic score informed by genome-wide association studies of multiple ancestries and related traits improves risk prediction for coronary artery disease 91%
- Genome-wide polygenic score with APOL1 risk genotypes predicts chronic kidney disease across major continental ancestries 90%
Similar papers in this journal
- Identification of Digital Twins to Guide Interpretable AI for Diagnosis and Prognosis in Heart Failure 95%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 93%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 93%
Similar papers in this journal
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 94%
- Clinical and genetic associations of asymmetric apical and septal left ventricular hypertrophy 93%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 92%
Similar papers in this journal
Similar papers in this journal
- Genome-wide association analysis and Mendelian randomization proteomics identify novel protein biomarkers and drug targets for primary prevention of heart failure 93%
- Deep representation learning for clustering longitudinal survival data from electronic health records 93%
- Weakly supervised classification of rare aortic valve malformations using unlabeled cardiac MRI sequences 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.