Reward-Guided Generation Improves the Scientific Utility of Synthetic Biomedical Data
Jackson, N. J.; Espinosa-Dice, N.; Yan, C.; Malin, B. A.
Show abstract
Synthetic data generation is a promising approach for biomedical data sharing and dataset augmentation, yet existing methods lack mechanisms to preserve statistical properties necessary for scientific analysis. To address this, we introduce RLSYN+REG, a reinforcement learning-driven generative model, which encourages that regression models trained on synthetic data reproduce the coefficients and predictions of their real-data counterparts. We evaluate RL-SO_SCPLOWYNC_SCPLOW+RO_SCPLOWEGC_SCPLOW on MIMIC-III and the American Community Survey (ACS) across regression model reproduction, fidelity to real data, and privacy. Synthetic data from RLSO_SCPLOWYNC_SCPLOW+RO_SCPLOWEGC_SCPLOW substantially improves upon that of RLSO_SCPLOWYNC_SCPLOW, raising correlations between real and synthetic regression coefficients from 0.054 to 0.600 on MIMIC-III and from 0.160 to 0.376 on ACS. Predictive performance also improves, reducing the gap between real-data baselines by 81.4% and 97.6% on MIMIC-III and ACS, respectively. These improvements come with negligible cost to fidelity or privacy and are robust to reductions in training data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 96%
- Privacy-Preserving Federated Neural Network Learning for Disease-Associated Cell Classification 96%
- Tokenized and Continuous Embedding Compressions of Protein Sequence and Structure 94%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 95%
- GRouNdGAN: GRN-guided simulation of single-cell RNA-seq data using causal generative adversarial networks 94%
- Deep transfer learning for reducing health care disparities arising from biomedical data inequality 93%
Similar papers in this journal
- Actionability of Synthetic Data in a Heterogeneous and Rare Healthcare Demographic; Adolescents and Young Adults (AYAs) with Cancer 95%
- A Bayesian Framework for Detecting Gene Expression Outliers in Individual Samples 91%
- Error reduction in leukemia machine learning classification with conformal prediction 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.