Back

Foundation model embeddings enable cardiovascular screening for people living with HIV in Vietnam using wearable signals

Mesinovic, M.; Bich, H. H.; Trieu, L. V.; Quoc, V. N.; Thanh, N. N.; Hoang, T. A. N.; Hoang, M. T. V.; Khanh, P. N. Q.; Vo, X. H.; Hong, P. V.; Van, K. L. D.; Minh, Y. L.; Minh, Y. L.; Thwaites, L.; Zhu, T.

2025-10-13 health informatics
10.1101/2025.10.11.25337800 medRxiv
Show abstract

Cardiovascular disease (CVD) screening faces significant challenges in resource-limited settings, where infrastructure and computational constraints preclude the use of advanced remote assessment. These constraints are particularly acute for people living with HIV (PLWH), who experience elevated CVD risk yet often receive care in clinics without the capacity for specialist diagnostics. We evaluate pretrained physiological embeddings from foundation models for CVD detection using low-cost wearable photoplethysmography (PPG) signals from 80 PLWH outpatients in Ho Chi Minh City, Vietnam. We compare a strictly zero-shot approach (NormWear applied without any local training) with a more practical pipeline that uses frozen PaPaGei embeddings plus a locally trained classifier. The PaPaGei-embedding approach achieved superior discrimination (AUROC 0.769) compared with zero-shot NormWear (0.610), traditional PCA features (0.651), and established clinical scores, including the Framingham score (0.551) and the D:A:D-modified Framingham score (0.462). Without fine-tuning the foundation model itself, PaPaGei embeddings captured clinically coherent structure: patients on dolutegravir-based regimens clustered in low-risk regions, while those with high cholesterol variability occupied high-risk areas, consistent with cardiometabolic pathophysiology. These results show that pretrained physiological embeddings can enable accurate screening when combined with lightweight local calibration, reducing reliance on extensive feature engineering while preserving pathophysiological plausibility and actionable triage behaviour in the learned embeddings. This provides a practical approach for deploying foundation models in resource-constrained settings where training deep models may be infeasible.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.