使用分布式医疗补助数据进行异质治疗效果的分布式融合R学习
Jinhong Li1, Julie M Donohue2, Lu Tang1
1Department of Biostatistics and Health Data Science, University of Pittsburgh, Pittsburgh, PA 15261, United States.
Biometrics
|March 5, 2026
概括
本研究引入了一种保护隐私的分布式融合学习方法 (DFR-learner),用于在不共享敏感参与者数据的情况下,在多个数据站点中估计异质治疗效应.
科学领域:
- 医疗信息学 医疗信息学
- 统计学学习 统计学学习
- 保护隐私的数据分析数据分析
背景情况:
- 数据驱动的决策需要对异质治疗效应进行准确的估计.
- 整合来自多个站点的数据可以改善样本大小,以便进行可靠的CATE估计.
- 挑战包括跨站点的治疗效果异质性和数据隐私问题.
研究的目的:
- 开发一种方法,在分布式数据站点上共同估计条件平均治疗效应 (CATE).
- 在数据集成中应对治疗效应异质性和隐私保护方面的挑战.
- 为了实现高效和私有信息交换,改进CATE估计.
主要方法:
- 提出了一个分布式融合学习方法,DF R-learner.
- 在不合并个人参与者数据的情况下,共同估计CATE跨站点的估计.
- 采用数据驱动的融合惩罚,以结合相似的参数和信任分布,以实现高效的私人交换.
主要成果:
- DF R-learner允许在不同站点上使用不同的CATE功能.
- 通过结合相似的参数来实现改进的估计.
- 理论上和经验上证明,与集中数据方法相比,效率没有损失.
- 通过使用分布式医疗补助数据,成功应用于研究阿片类药物使用障碍的药物治疗.
结论:
- DF R-learner有效地以分布式,保护隐私的方式估计CATE.
- 该方法解决了现实世界数据集成的关键挑战.
- 提供了一种可行的解决方案,可以利用多个站点的数据,同时保护敏感信息.
相关概念视频
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
301
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
301
Distributions to Estimate Population Parameter
5.3K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.3K
Bioequivalence Experimental Study Designs: Repeated Measures, Cross-Over, Carry-Over, and Latin Square Designs
323
Bioequivalence experimental study designs play a pivotal role in testing the effectiveness of various treatments. Key among these are the repeated measures, cross-over, carry-over, and Latin square designs. In the repeated measures design, each subject receives all treatments, allowing for temporal comparisons. This type of design is useful in reducing variability but requires careful planning to avoid bias.The cross-over design, an economical method, involves sequential administration of...
323
Friedman Two-way Analysis of Variance by Ranks
523
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
523
Randomized Experiments
9.2K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
9.2K


