将倾向得分方法与变异自编码器结合起来,在存在潜伏子组时生成合成数据
Kiana Farhadyar1,2, Federico Bonofiglio3, Maren Hackenberg4,5
1Institute of Medical Biometry and Statistics, University of Freiburg, Freiburg, Germany. kiana.farhadyar@uniklinik-freiburg.de.
BMC medical research methodology
|September 9, 2024
概括
生成合成临床数据需要保持个体差异. 这项研究将变量自编码器与统计方法相结合,以准确地反映在合成数据集中的已知和未知的患者异质性.
科学领域:
- * 计算统计学统计学
- * 机器学习 * 机器学习
- * 生物信息学是一门学科.
背景情况:
- * 临床数据隐私法规需要合成数据生成.
- * 临床队列的异质性,无论是已知的 (子组) 还是未知的 (分布性质),都给合成数据带来了挑战.
- *变量自编码器 (VAE) 是用于生成合成数据的深度学习模型,但可能会与复杂的异质性作斗争.
研究的目的:
- * 调查保存和控制VAE产生的合成数据中已知和未知的异质性的方法.
- * 开发一种方法,忠实地复制复杂的分布性质和临床队列的子组特征.
- *使用模拟和真实世界的临床试验数据来评估拟议的方法.
主要方法:
- *将变量自编码器 (VAE) 与预转换相结合,以捕捉边际分布中的未知异质性.
- *将VAE与倾向性得分回归模型集成,以管理来自子组成员身份的已知异质性.
- *使用现实的模拟设计和来自国际中风试验的真实数据来评估该方法.
主要成果:
- * 提出的以VAE为基础的方法与预转换成功地复制了具有挑战性的边际分布,优于仅关注边际分布的方法.
- * 倾向分数提供了有价值的补充信息,有助于隐性空间的可视化,并允许对具有或没有子组特征的合成数据进行受控抽样.
- * 该方法有效地处理现实世界的临床数据,显示特定地点的分布差异和双模式.
结论:
- *将生成深度学习 (VAE) 与倾向得分回归等统计方法相结合,对于生成准确反映队列异质性的合成临床数据至关重要.
- * 拟议的方法提供了一个强大的框架,用于控制和保护合成数据集中已知和未知变异源.
- * 这种方法提高了合成数据的研究和开发的实用性,同时遵守数据保护法规.
相关概念视频
Variability: Analysis
133
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
133
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
423
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
423
Stratified Sampling Method
11.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
11.9K
Randomized Experiments
6.8K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.8K
Variance
9.3K
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.
The standard deviation measures the spread in the same units as the...
The standard deviation measures the spread in the same units as the...
9.3K
Multi-input and Multi-variable systems
103
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
103


