Validating methods for inferring co-occurring diseases: a flexible framework for simulating synthetic data
Marchi, H.; Schmiegel, S.; Fuchs, C.; Schamberger, T.
Show abstract
The validation of methods is an integral part of statistical research, defining conditions under which methods can yield reliable results. When validation is carried out empirically, it requires a solid data basis that allows the control and management of relevant characteristics. In particular, depending on the context, the data must meet specific requirements regarding, e. g. sample size, dimensionality, completeness and underlying dependency structures. Real-world data often fails to meet these requirements, particularly in medical contexts where availability or the right to publish is additionally restricted through privacy regulations. For this reason, synthetic data is an effective alternative for method validation. Generating synthetic data is particularly demanding if it is required to precisely mirror complex dependence structures while simultaneously controlling certain target characteristics. In our work, we address the medical context of simultaneously co-occurring diseases, where symptoms may overlap or conflict. We seek to generate synthetic data supporting the simulation-based validation of statistical methods which are able to predict holistic disease pictures based on patient information, including the accounting for comorbidities. We introduce a four-step framework in which we (I) generate patient covariates such as symptoms; (II) connect this patient information to predictors for the single or joint occurrence of diseases; (III) transform the predictors into disease probabilities or disease scores; (IV) convert the probabilities or scores into disease occurrence. Within each of these steps, we outline several alternatives that allow different forms of modeling the overall dependence structure. We apply our framework to a case study informed by real-world data: to the context of pain-causing diseases which share certain similarities in their clinical presentations, and which can occur either individually or jointly. In this analysis, we employ five combinations of methodological alternatives within the data generation steps. This allows us to evaluate the data generation approaches with respect to their ability to achieve the defined target characteristics, and to demonstrate strengths and weaknesses as well as specific suitability. The proposed simulation framework is broadly applicable beyond the specific use case and the medical context. The approaches are designed to be accessible and adjustable through varying input settings, enabling users to tailor the data generation to their specific needs. This way, our work provides researchers with a flexible framework for generating synthetic validation data that aligns with the methodological requirements of their studies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Tree-informed Bayesian multi-source domain adaptation: cross-population probabilistic cause-of-death assignment using verbal autopsy 96%
- A scalable approach for continuous time Markov models with covariates 96%
- Survival Analysis on Rare Events Using Group-Regularized Multi-Response Cox Regression 95%
Similar papers in this journal
Similar papers in this journal
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 96%
- Individual Reference Intervals for Personalized Interpretation of Clinical and Metabolomics Measurements 95%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 95%
Similar papers in this journal
- Probabilistic Cause-of-disease Assignment using Case-control Diagnostic Tests: A Latent Variable Regression Approach 96%
- A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Québec Administrative Data 95%
- Bias reduction and inference for electronic health record data under selection and phenotype misclassification: three case studies 95%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 94%
- Using indication embeddings to represent patient health for drug safety studies 93%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.