Back

Longitudinal Clinical Foundation Models Augmented with Genomics for Early Detection and Risk Stratification of Inherited Cardiomyopathy

Zolensky, A. L.; Kripke, C. M.; Keat, K.; Damrauer, S. M.; Levin, M. G.; Verma, A.

2026-08-12 health informatics
10.64898/2026.08.10.26360107 medRxiv
Show abstract

Hypertrophic and dilated cardiomyopathy (HCM and DCM) carry substantial morbidity and mortality, yet diagnosis may be delayed, particularly when presentation is nonspecific. Existing machine-learning approaches to cardiomyopathy phenotyping, genotype prediction, and risk stratification commonly rely on disease-specific, hand-engineered features drawn from echocardiography, cardiac MRI, ECG, or curated clinical variables. We evaluated whether a general-purpose clinical foundation model, CLMBR-T-base, pre-trained via next-clinical-event prediction with no cardiomyopathy-specific supervision, could produce linearly separable embeddings for all three case/control cohorts. Using EHR data from the Penn Medicine BioBank, we constructed cohorts for (1) prediction of a first recorded qualifying HCM/DCM diagnosis at 1-, 3-, and 6-month horizons, decomposed into eventual-versus-never-case and imminent-versus-eventual comparisons; (2) genetic carrier status prediction among diagnosed patients with completed gene panels; and (3) prediction of heart-failure hospitalization, and all-cause mortality as both binary and time-to-event outcomes. Linear probes fitted to frozen embeddings achieved AUROCs of 0.75-0.82 for onset prediction, 0.74-0.75 for genotype status, and Harrell's concordance of 0.65-0.80 for time-to-event outcomes. Decomposing the onset prediction task reveals that the model often misclassifies patients who were diagnosed later as positive, suggesting the patient journey embeddings encode disease state more reliably than care timing. These results suggest that a single, generically pretrained EHR embedding can support multiple clinically motivated prediction problems in CM without disease-specific feature engineering.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.