生成和评估合成纵向患者数据的方法:系统性审查
Katariina Perkonoja1,2, Kari Auranen1,3, Joni Virta1
1Department of Mathematics and Statistics, University of Turku, Turku, Finland.
Journal of healthcare informatics research
|February 9, 2026
概括
生成合成纵向患者数据对于医疗保健研究至关重要,但目前的方法缺乏隐私保护和明确的评估标准. 为了在现实世界中应用,需要进一步开发.
科学领域:
- 医疗信息学 医疗信息学
- 数据科学数据科学数据科学
- 医学研究 医学研究
背景情况:
- 由于隐私和安全问题,医疗保健数据的使用受到阻碍.
- 合成数据生成为敏感的患者信息提供了一种保护隐私的替代方案.
- 纵向患者数据方面在现有的合成数据审查中表现不足.
研究的目的:
- 系统地审查生成和评估合成纵向患者数据的方法.
- 识别当前合成纵向患者数据生成技术中的挑战和差距.
主要方法:
- 按照PRISMA指南进行系统的文献审查.
- 搜索了5个数据库,截至2024年5月.
- 确定和分析了39种合成纵向患者数据生成方法.
主要成果:
- 四种方法解决了关键挑战:时间结构,变量类型,缺失值和数据不平衡.
- 大多数研究评估了数据的相似性和实用性;较少的研究评估了隐私.
- 没有确定的方法集成隐私保护机制,小数据集的有效性尚不清楚.
结论:
- 目前的合成纵向患者数据方法需要加强隐私和强大的评估.
- 标准化的评估标准和隐私保护机制对于现实世界的适用性至关重要.
- 未来的研究应该侧重于隐私,评估框架,可访问的代码和监管指导.
相关概念视频
Systematic Sampling Method
13.4K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
Systematic sampling is one of the simplest methods...
Systematic sampling is one of the simplest methods...
13.4K
Longitudinal Research
13.4K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
13.4K
Review and Preview
8.4K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
8.4K
Review and Preview
11.6K
Data are individual items of information obtained from a population or sample. Data may be classified as qualitative (categorical), quantitative continuous, or quantitative discrete. Because it is not practical to measure the entire population in a study, researchers use samples to represent the population. A random sample is a representative group from the population chosen by using a method that gives each individual in the population an equal chance of being included in the sample. Random...
11.6K
Longitudinal Studies
535
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
535
Random and Systematic Errors
15.2K
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
15.2K


