Synthetic Longitudinal Tabular Data Generation via Copula
Cai, H.; Yu, W.; Lu, R.; Chattopadhyay, I.; Zhang, X.; Liu, J.
Show abstract
Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existing generative methods, particularly generative adversarial network (GAN)-based approaches, can model complex distributions, but their estimated dependence structures are often difficult to interpret and their performance may be unstable or prone to overfitting in modestly sized datasets. Here we show that eCDF-copula, a statistically rooted approach using the empirical cumulative distribution function (eCDF) and copula modeling, preserves within- and between-visit dependence structure. To handle pervasive missing data, we propose a two-stage strategy combining multiple imputation with copula-based synthesis, enabling a variance decomposition that quantifies replication variability across methods. We benchmarked the proposed approach against four established methods on two longitudinal clinical datasets spanning markedly different sample sizes (n = 120 vs. n = 3, 612). eCDF-copula achieved resemblance and utility exceeding those of state-of-the-art synthetic data methods, while maintaining comparable privacy.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generation of realistic synthetic data using multimodal neural ordinary differential equations 93%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 93%
- FedWeight: Mitigating Covariate Shift of Federated Learning on Electronic Health Records Data through Patients Re-weighting 92%
Similar papers in this journal
Similar papers in this journal
- sureLDA: A Multi-Disease Automated Phenotyping Method for the Electronic Health Record 93%
- ATLAS: An automated association test using probabilistically linked health records with application to genetic studies 91%
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.