A novel pipeline for realistic synthetic longitudinal EHR data generation
Gabrielle Josling1, Ibrahima Diouf1, Sankalp Khanna1
1Commonwealth Scientific and Industrial Research Organisation.
Research Square
|February 6, 2026
Summary
A new pipeline generates realistic synthetic electronic health record data, preserving key statistical properties for machine learning and statistical modeling. This synthetic data can augment or replace real data in many analytical workflows, enhancing privacy and data sharing capabilities.
Area of Science:
- Health Informatics
- Data Science
- Computational Biology
Background:
- Synthetic health data generation is crucial for patient privacy and data sharing.
- Existing methods often lack structural realism and are narrowly evaluated.
- This limits the practical application of synthetic data in downstream analytical workflows.
Purpose of the Study:
- Introduce a novel pipeline for generating realistic synthetic longitudinal electronic health record (EHR) data.
- Evaluate the pipeline's performance across three diverse datasets.
- Provide evidence-based guidance on using synthetic data to replace or augment real data.
Main Methods:
- Extended existing HALO and ConSequence frameworks with a post-processing step for continuous variables and timestamps.
- Applied the pipeline to small longitudinal, medium intensive-care, and large multi-hospital administrative datasets.
- Assessed realism and utility for machine learning, statistical modeling, and time series analysis.
Main Results:
- Generated realistic synthetic data preserving key statistical properties and relationships across all datasets.
- Machine learning models trained on synthetic data showed comparable predictive accuracy and feature importance to those trained on real data.
- Statistical modeling results closely aligned with real data, though precision for rare conditions may be limited; time series analysis was unsuitable.
Conclusions:
- The pipeline successfully produces realistic and analytically useful synthetic longitudinal EHR data across various scales.
- Synthetic data demonstrates strong utility for machine learning and statistical modeling tasks.
- Findings offer practical guidance for the judicious use of synthetic data in healthcare analytics.
More Related Videos
Related Concept Videos
Longitudinal Research
13.4K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
13.4K
Longitudinal Studies
533
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
533
Synthetic Biology
5.6K
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
5.6K
Synthetic Disvision of Polynomials
190
Synthetic division is an efficient algorithmic approach for dividing a polynomial by a linear binomial of the form x - c, where c is a real number. This method is helpful due to its streamlined process, which avoids the more cumbersome steps involved in the traditional long division of polynomials. It simplifies computation and serves as a practical tool for evaluating polynomials and identifying their factors.To perform synthetic division, one begins by listing the coefficients of the...
190
How Data are Classified: Categorical Data
44.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.8K
How Data are Classified: Numerical Data
38.1K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
38.1K


