安全的分布式多重归算使得私人数据所有者能够推断缺失的数据
Haris Smajlović1, Yi Lian2, Qi Long2
1Department of Computer Science, University of Victoria, Victoria, BC, Canada.
NPJ digital medicine
|January 9, 2026
概括
本研究介绍了一种使用安全多方计算 (SMC) 来协同分析私人电子健康记录 (EHR) 的安全方法. 这种方法使准确的数据归算成为可能,并改善了重症监护室 (ICU) 患者结果的分类.
科学领域:
- 医疗信息学 医疗信息学
- 计算安全计算安全
- 生物统计学 生物统计学
背景情况:
- 电子健康记录 (EHR) 对医学研究至关重要,但在各个机构中通常是分散的.
- 隐私问题和数据不完整性阻碍了使用分布式EHR的协作研究.
- 目前的方法很难将私人数据集中用于全面分析和归算.
研究的目的:
- 开发一种安全,保护隐私的解决方案,用于对分布式电子健康记录进行协作分析.
- 为了使不完整的EHR数据集中缺少数据的准确统计归算.
- 改善重症监护机构对高风险患者的分类.
主要方法:
- 使用安全多方计算 (SMC) 实现可证明安全的解决方案.
- 允许分布式数据集作为一个整体用于归算和集体研究.
- 在合成和现实数据集上测试解决方案.
主要成果:
- 基于SMC的解决方案实现了与非安全方法相比的实际运行时间和准确性.
- 这种方法有效地归因于分布式EHR中缺少的数据.
- 在ICU入院期间对高风险患者的分类结果显著改善.
结论:
- 安全的多方计算为保护隐私的协作EHR研究提供了可行的解决方案.
- 开发的方法克服了数据隐私和不完整性的局限性.
- 这有助于进行更全面的研究,并加强对患者结果的临床决策.
相关概念视频
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
447
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
447
Prediction Intervals
3.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.2K
Censoring Survival Data
518
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
518
Distributions to Estimate Population Parameter
5.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.0K
Assumptions of Survival Analysis
391
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
391
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K


