D3MI:一种高效,强大的联合归算方法,用于减少分布式不完整数据分析中的偏差,通过考虑网站内部的相关性和网站之间的异质性
Yi Lian1, Xiaoqian Jiang2, Qi Long1
1Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA, USA.
medRxiv : the preprint server for health sciences
|May 19, 2025
概括
我们开发了一种新的方法,即基于分布式混合模型的多重推断 (Distributed Mixed Model-based Multiple Imputation,D3MI),用于在多个机构的电子健康记录 (EHR) 中准确处理缺失的数据,同时保持患者的隐私.
科学领域:
- 医疗信息学 医疗信息学
- 生物统计学 生物统计学
- 机器学习 机器学习
背景情况:
- 电子健康记录 (EHR) 对临床研究有价值,但含有缺失的数据,引入了偏见.
- 现有的分布式归算方法在EHR数据中的站内相关性和站间变异性方面存在困难.
- 隐私问题限制了分布式EHR数据集的直接分析.
研究的目的:
- 引入一种新的联合归算方法,即基于分布式混合模型的多重归算 (Distributed Mixed Model-based Multiple Imputation,简称D3MI),以解决分布式EHR中缺少数据的挑战.
- 开发一种方法,以考虑EHR数据中的网站内部相关性和网站间异质性.
- 增强分布式临床数据的隐私保护分析.
主要方法:
- D3MI集成了联合学习,相关数据的统计方法和多层次归算.
- 它明确地使用特定地点的随机效应来模拟网站内部的相关性和网站之间的异质性.
- 该方法通过避免原始数据共享来确保隐私,并且具有计算效率.
主要成果:
- 模拟显示D3MI在准确性和一致性方面优于最先进的分布式归算方法.
- D3MI成功地应用于来自格鲁吉亚科弗德尔急性中风注册表的现实世界EHR病例研究.
- 该方法在分布式环境中有效处理不完整和聚类数据.
结论:
- D3MI为分布式,对隐私敏感的EHR设置中缺少数据提供了强大的解决方案.
- 通过建模复杂的数据结构,D3MI提高了协作临床研究的严谨性和可重复性.
- 这种方法增强了研究目的的联合EHR数据的实用性.
更多相关视频
相关概念视频
Bias in Epidemiological Studies
120
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
120
Friedman Two-way Analysis of Variance by Ranks
130
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
130
Distributions to Estimate Population Parameter
4.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.0K
Stratified Sampling Method
11.7K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
11.7K
Bonferroni Test
2.7K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.7K
Bias
3.7K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.7K


