Synthetic Health Data Can Augment Community Research Efforts to Better Inform the Public During Emerging Pandemics
Prasanna, A.; Jing, B.; Plopper, G.; Krasnov Miller, K.; Sanjak, J.; Feng, A.; Prezek, S.; Vidyaprakash, E.; Thovarai, V.; Maier, E.; Bhattacharya, A.; Naaman, L.; Stephens, H.; Watford, S.; Boscardin, W. J.; Johanson, E.; Lienau, A.
Show abstract
The COVID-19 pandemic had disproportionate effects on the Veteran population due to the increased prevalence of medical and environmental risk factors. Synthetic electronic health record (EHR) data can help meet the acute need for Veteran population-specific predictive modeling efforts by avoiding the strict barriers to access, currently present within Veteran Health Administration (VHA) datasets. The U.S. Food and Drug Administration (FDA) and the VHA launched the precisionFDA COVID-19 Risk Factor Modeling Challenge to develop COVID-19 diagnostic and prognostic models; identify Veteran population-specific risk factors; and test the usefulness of synthetic data as a substitute for real data. The use of synthetic data boosted challenge participation by providing a dataset that was accessible to all competitors. Models trained on synthetic data showed similar but systematically inflated model performance metrics to those trained on real data. The important risk factors identified in the synthetic data largely overlapped with those identified from the real data, and both sets of risk factors were validated in the literature. Tradeoffs exist between synthetic data generation approaches based on whether a real EHR dataset is required as input. Synthetic data generated directly from real EHR input will more closely align with the characteristics of the relevant cohort. This work shows that synthetic EHR data will have practical value to the Veterans health research community for the foreseeable future.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 95%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 94%
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 93%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 95%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 94%
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 93%
Similar papers in this journal
- Assessing the effects of data drift on the performance of machine learning models used in clinical sepsis prediction 94%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 94%
- Image and structured data analysis for prognostication of health outcomes in patients presenting to the Emergency Department during the COVID-19 pandemic 93%
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 94%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 94%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.