High-Fidelity Synthetic Data Replicates Clinical Prediction Performance in a Million-Patient Diabetes Cohort
de la Oliva-Roque, V. M.; Kreil, D. P.; Dopazo, J.; Ortuno, F.; Loucera, C.
Show abstract
Synthetic data generated using generative models trained on real clinical data offers a promising solution to privacy concerns in health research. However, many efforts are limited by small or demographically narrow training datasets, reducing the generalizability of the synthetic data. To address this, we used real-world clinical data from nearly one million individuals with diabetes in the Andalusian Population Health Database (BPS) to generate a comprehensive longitudinal synthetic dataset. We employed a dual adversarial autoencoder to produce synthetic data and evaluated its utility in a clinical machine learning (ML) task: predicting the onset of chronic kidney disease, a common diabetes complication. Models trained on synthetic data were assessed for their ability to reproduce patterns and predictive behaviors observed in real data. Performance and stability were compared across models trained on real, synthetic, and hybrid datasets. Models trained exclusively on synthetic data achieved AUROC scores comparable to real-data models (0.70 vs. 0.73) and showed high stability in feature importance rankings (weighted Kendalls {tau} > 0.9). Notably, combining synthetic and real data did not improve performance. Our findings demonstrate that high-fidelity synthetic longitudinal data can replicate real data performance in clinical ML, supporting its use in research while preserving patient privacy. This represents a significant step toward more collaborative and privacy-preserving healthcare data ecosystems.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 94%
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 94%
- Modeling physician variability to prioritize relevant medical record information 93%
Similar papers in this journal
Similar papers in this journal
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 92%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 92%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 91%
Similar papers in this journal
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 94%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- High-throughput Phenotyping with Temporal Sequences 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.