o-HD:分布式因果推理与共变量转移用于分析现实世界高维数据.
Jiayi Tong1, Jie Hu1, George Hripcsak2
1Department of Biostatistics, Epidemiology and Informatics, University of Pennsylvania, Philadelphia, PA 19104, USA.
概括
这项研究介绍了DisC2o-HD,这是一种分布式学习算法,用于高维的医疗数据. 它有效地估计了平均治疗效应 (ATE),同时解决了多个临床场所的共变量转移.
科学领域:
- 医疗信息学 医疗信息学
- 生物统计学 生物统计学
- 机器学习 机器学习
背景情况:
- 高维的医疗保健数据 (EHR,索赔) 带来了挑战:许多变量,多站点数据整合和共同变量转移.
- 在此类数据中估计治疗效应需要强大的方法来处理跨站点的异质性.
研究的目的:
- 提出一种新的分布式学习算法DisC2o-HD,用于在高维的医疗数据中估计平均治疗效果 (ATE).
- 针对多个临床场所的共同变量转移和数据异质性.
主要方法:
- 开发了DisC2o-HD,这是一个利用替代概率的分布式学习算法.
- 采用倾向分数和结果模型校准来实现共变量平衡,并考虑共变量转移.
- 证明分布式估计器接近聚合估计器.
主要成果:
- 如果正确指定倾向性得分或结果回归模型,则建议的估计器是一致的.
- 当两个模型都被正确指定时,实现半参数效率.
- 模拟研究和现实世界的数据应用验证了算法的性能和准备.
结论:
- DisC2o-HD提供了一种有效和可实施的解决方案,用于在分布式,高维的医疗保健数据中估计ATE.
- 该算法提供了一种强大的方法,可以利用多站点数据,同时保持统计有效性.
相关概念视频
Causality in Epidemiology
863
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
863
Variability: Analysis
191
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
191
Friedman Two-way Analysis of Variance by Ranks
299
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
299
Correlation and Causation
39.6K
Statistical tests can calculate whether there is a relationship, or correlation, between independent and dependent variables. An indirect relationship of the variables signifies a correlation, while a direct relationship shows causation. If it is determined that no connection exists between the variables, then the correlation is a coincidence.
Correlation versus Causation
If the dependent variable increases or decreases when the independent variable increases, there is a positive or negative...
Correlation versus Causation
If the dependent variable increases or decreases when the independent variable increases, there is a positive or negative...
39.6K
Statistical Methods for Analyzing Epidemiological Data
539
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
539
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
215
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
215


