Related Experiment Video
Updated: Aug 21, 2026

In Vivo Modeling of the Morbid Human Genome using Danio rerio
Published on: August 24, 2013
Validating methods for inferring co-occurring diseases: a flexible framework for simulating synthetic data
Hannah Marchi1, Sophie Schmiegel1, Tamara Schamberger1
1Data Science Group, Faculty of Business Administration and Economics, Bielefeld University, Bielefeld, Germany.
Background:
The validation of methods is an integral part of statistical research, defining conditions under which methods yield reliable results. Empirical validation requires a solid data basis to control and manage relevant characteristics like sample size, dimensionality, and underlying dependency structures. Real-world data often fails to meet these requirements, particularly in medical contexts where privacy regulations restrict availability. For this reason, synthetic data is an effective alternative for method validation. However, generating synthetic data is demanding when it must precisely mirror complex dependence structures while simultaneously controlling specific target characteristics.
Methods:
We address the medical context of co-occurring diseases, where symptoms may overlap or conflict. We propose a four-step framework to generate synthetic data for the simulation-based validation of statistical methods. The framework involves: (I) generating patient covariates; (II) connecting this information to predictors for single or joint disease occurrence; (III) transforming predictors into disease probabilities or scores; and (IV) converting these into disease occurrences. Each step offers several alternatives for modeling the overall dependence structure. We apply our framework to a case study of pain-causing diseases which share certain similarities in their clinical presentations, and which can occur either individually or jointly. By employing five combinations of methodological alternatives, we evaluate the approaches' ability to achieve target characteristics and demonstrate their specific strengths and weaknesses.
Results:
Matching the data-generating process with the estimation method allows for the successful recovery of input information, such as coefficients and correlations. Target properties like disease prevalence and associations are achieved to varying degrees depending on the methods used.
Conclusions:
While the proposed theory-driven framework is broadly applicable beyond the specific medical use case, it relies on careful, domain-informed parameter curation to generate meaningful synthetic datasets. Its flexible, adjustable input settings enable researchers to tailor data generation to their precise methodological requirements, providing a controlled basis for simulation-based validation without implying direct clinical inference.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe and...
Principles of Disease Surveillance
Statistical Methods for Analyzing Epidemiological Data
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Investigation of Disease Outbreaks
Statistical Software for Data Analysis and Clinical Trials