个人数据受保护 高维异构数据的整合回归分析
Tianxi Cai1, Molei Liu1, Yin Xia2
1Department of Biostatistics, Harvard School of Public Health, Harvard University, Boston, USA.
Journal of the American Statistical Association
|November 17, 2023
概括
我们开发了SHIR,一种用于整合多个研究中的高维数据而无需共享单个数据的新方法. 通过SHIR,即使数据异质,也可以进行一致的变量选择和高效估计.
科学领域:
- 生物统计学 生物统计学
- 计算生物学 计算生物学
- 医疗信息学 医疗信息学
背景情况:
- 超分析提高了研究的精度和通用性,但在超高维数据方面面临着挑战.
- 整合异质研究是复杂的,特别是在DataSHIELD限制下,防止个人共享数据.
研究的目的:
- 提出一种新的整合估计程序,SHIR (数据屏蔽高维整合回归),用于具有DataSHIELD约束的超高维设置.
- 开发一种方法,保护个人数据,同时适应研究间的异质性,并允许一致的变量选择.
主要方法:
- SHIR使用基于总结统计的整合程序来保护个人数据隐私.
- 该方法适应跨研究的共变量分布和稀疏回归模型参数的异质性.
- 在理论上,SHIR与现有的分布式方法相比较,显示出优越的统计效率.
主要成果:
- 在超高维设置中,SHIR实现了一致的变量选择.
- 总结总结统计数据的估计误差是可以忽略不计的,接近统计最小值率.
- SHIR在异面上相当于所有数据共享的理想估计器.
结论:
- SHIR提供了一个统计学上高效且保护隐私的解决方案,用于在多项研究中对高维数据进行综合分析.
- 与现有的分布式方法相比,该方法显示出更高的性能.
- 通过应用到冠状动脉疾病表型的电子健康记录来验证SHIR的实用性.
相关概念视频
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Statistical Analysis: Overview
6.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
6.6K
Friedman Two-way Analysis of Variance by Ranks
206
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
206
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Mechanistic Models: Compartment Models in Individual and Population Analysis
43
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
43


