Generating Synthetic Multi-national Longitudinal Cohorts for Clinically Grounded HIV Research
Liang, Z. J.; Li, Z.; Jackson, N. J.; Caro-Vega, Y.; Moreira, R. I.; Paredes, F.; Bernadin, J.; Varela, D.; Cesar, C.; Blasimme, A.; Perkins, J. M.; Asiaee, A.; Duda, S. N.; Malin, B. A.; Shepherd, B. E.; Yan, C.
Show abstract
High-quality, widely accessible international longitudinal cohort data for people living with HIV (PWH) have long been needed for advancing open science and data-driven innovation, yet stringent and incongruent privacy regulations have made data sharing difficult. Synthetic data generation offers a promising privacy-preserving alternative, but producing realistic synthetic cohorts of PWH remains challenging due to complex temporal dynamics, interdependent clinical variables, long follow-up periods, and high missingness inherent in such data. Here, we introduce Medical Longitudinal latent Diffusion (MeLD), a generative model designed to synthesize variable-length, decades-spanning, mixed-type clinical trajectories with missingness. Using the Caribbean, Central, and South America Network for HIV Epidemiology (CCASAnet) cohort, one of the worlds largest international HIV datasets with over 30 years of follow-up on nearly 50,000 PWH, we show that MeLD consistently outperforms state-of-the-art methods across data utility, fidelity, and privacy. Notably, MeLD excels in longitudinal inference utility, accurately reproducing time-to-death estimates and risk factor effects, while maintaining strong privacy protection. This work delivers the first in-depth, large-scale, and openly accessible synthetic longitudinal cohort of PWH that faithfully preserves the distributional patterns and clinical associations observed in real data, offering an immediately deployable resource for hypothesis generation, methods innovation, medical training, and reproducible HIV research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- FedWeight: Mitigating Covariate Shift of Federated Learning on Electronic Health Records Data through Patients Re-weighting 96%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 95%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Predicting clinical drug response from model systems by non-linear subspace-based transfer learning 95%
- Robust probabilistic modeling for single-cell multimodal mosaic integration and imputation via scVAEIT 94%
- Dissecting heterogeneous cell-populations across drug and disease conditions with PopAlign 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.