Back

Validating methods for inferring co-occurring diseases: a flexible framework for simulating synthetic data

Marchi, H.; Schmiegel, S.; Fuchs, C.; Schamberger, T.

2026-01-15 health informatics
10.64898/2026.01.13.26344018 medRxiv
Show abstract

The validation of methods is an integral part of statistical research, defining conditions under which methods can yield reliable results. When validation is carried out empirically, it requires a solid data basis that allows the control and management of relevant characteristics. In particular, depending on the context, the data must meet specific requirements regarding, e. g. sample size, dimensionality, completeness and underlying dependency structures. Real-world data often fails to meet these requirements, particularly in medical contexts where availability or the right to publish is additionally restricted through privacy regulations. For this reason, synthetic data is an effective alternative for method validation. Generating synthetic data is particularly demanding if it is required to precisely mirror complex dependence structures while simultaneously controlling certain target characteristics. In our work, we address the medical context of simultaneously co-occurring diseases, where symptoms may overlap or conflict. We seek to generate synthetic data supporting the simulation-based validation of statistical methods which are able to predict holistic disease pictures based on patient information, including the accounting for comorbidities. We introduce a four-step framework in which we (I) generate patient covariates such as symptoms; (II) connect this patient information to predictors for the single or joint occurrence of diseases; (III) transform the predictors into disease probabilities or disease scores; (IV) convert the probabilities or scores into disease occurrence. Within each of these steps, we outline several alternatives that allow different forms of modeling the overall dependence structure. We apply our framework to a case study informed by real-world data: to the context of pain-causing diseases which share certain similarities in their clinical presentations, and which can occur either individually or jointly. In this analysis, we employ five combinations of methodological alternatives within the data generation steps. This allows us to evaluate the data generation approaches with respect to their ability to achieve the defined target characteristics, and to demonstrate strengths and weaknesses as well as specific suitability. The proposed simulation framework is broadly applicable beyond the specific use case and the medical context. The approaches are designed to be accessible and adjustable through varying input settings, enabling users to tailor the data generation to their specific needs. This way, our work provides researchers with a flexible framework for generating synthetic validation data that aligns with the methodological requirements of their studies.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.