Back

Synthetic Health Data Can Augment Community Research Efforts to Better Inform the Public During Emerging Pandemics

Prasanna, A.; Jing, B.; Plopper, G.; Krasnov Miller, K.; Sanjak, J.; Feng, A.; Prezek, S.; Vidyaprakash, E.; Thovarai, V.; Maier, E.; Bhattacharya, A.; Naaman, L.; Stephens, H.; Watford, S.; Boscardin, W. J.; Johanson, E.; Lienau, A.

2023-12-13 health informatics
10.1101/2023.12.11.23298687 medRxiv
Show abstract

The COVID-19 pandemic had disproportionate effects on the Veteran population due to the increased prevalence of medical and environmental risk factors. Synthetic electronic health record (EHR) data can help meet the acute need for Veteran population-specific predictive modeling efforts by avoiding the strict barriers to access, currently present within Veteran Health Administration (VHA) datasets. The U.S. Food and Drug Administration (FDA) and the VHA launched the precisionFDA COVID-19 Risk Factor Modeling Challenge to develop COVID-19 diagnostic and prognostic models; identify Veteran population-specific risk factors; and test the usefulness of synthetic data as a substitute for real data. The use of synthetic data boosted challenge participation by providing a dataset that was accessible to all competitors. Models trained on synthetic data showed similar but systematically inflated model performance metrics to those trained on real data. The important risk factors identified in the synthetic data largely overlapped with those identified from the real data, and both sets of risk factors were validated in the literature. Tradeoffs exist between synthetic data generation approaches based on whether a real EHR dataset is required as input. Synthetic data generated directly from real EHR input will more closely align with the characteristics of the relevant cohort. This work shows that synthetic EHR data will have practical value to the Veterans health research community for the foreseeable future.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
JAMIA Open
37 papers in training set
Top 0.1%
36.9%
2
JMIR Public Health and Surveillance
45 papers in training set
Top 0.1%
12.1%
3
npj Digital Medicine
97 papers in training set
Top 0.7%
6.7%
50% of probability mass above
4
Journal of Biomedical Informatics
45 papers in training set
Top 0.3%
6.2%
5
BMC Medical Informatics and Decision Making
39 papers in training set
Top 0.6%
4.7%
6
Journal of the American Medical Informatics Association
61 papers in training set
Top 0.7%
3.9%
7
JMIR Medical Informatics
17 papers in training set
Top 0.4%
3.5%
8
PLOS ONE
4510 papers in training set
Top 44%
2.7%
9
IEEE Journal of Biomedical and Health Informatics
34 papers in training set
Top 0.7%
2.5%
10
Scientific Reports
3102 papers in training set
Top 54%
1.8%
11
Patterns
70 papers in training set
Top 1%
1.6%
12
Scientific Data
174 papers in training set
Top 2%
1.2%
13
International Journal of Medical Informatics
25 papers in training set
Top 1%
0.9%
14
Frontiers in Public Health
140 papers in training set
Top 7%
0.9%
15
Journal of Medical Internet Research
85 papers in training set
Top 4%
0.8%
16
Computer Methods and Programs in Biomedicine
27 papers in training set
Top 0.9%
0.8%
17
PLOS Computational Biology
1633 papers in training set
Top 25%
0.7%
18
GigaScience
172 papers in training set
Top 3%
0.7%
19
Heliyon
146 papers in training set
Top 7%
0.7%
20
Journal of Personalized Medicine
28 papers in training set
Top 1%
0.7%
21
Epidemiology
26 papers in training set
Top 0.7%
0.6%
22
Frontiers in Digital Health
20 papers in training set
Top 2%
0.6%
23
Proceedings of the National Academy of Sciences
2130 papers in training set
Top 48%
0.6%