PS-SAM:在动态和高效地从历史数据中借用信息之前,倾向性得分集成的自适应混合物
Yuansong Zhao1, Peng Yang2,3, Glen Laird4
1Department of Biostatistics and Data Science, University of Texas Health Science Center at Houston, Houston, TX, USA.
Journal of biopharmaceutical statistics
|April 17, 2025
概括
这项研究引入了一种新方法,即倾向性得分集成自适应混合物 (PS-SAM) 的先验,以更好地利用随机对照试验 (RCT) 中的历史数据. 这种方法减少了未测量因素的偏差,改善了治疗效果估计.
科学领域:
- 生物统计学 生物统计学
- 临床试验方法论 临床试验方法论
- 健康 数据科学 数据科学
背景情况:
- 历史数据可以提高随机对照试验 (RCT) 的效率,并减少样本大小要求.
- 历史和当前试验数据之间的患者特征差异是一个挑战.
- 倾向性得分方法 (匹配,反向概率加权) 根据基线异质性进行调整,但易受未测量的混因素的影响.
研究的目的:
- 开发一个强大的统计方法来将历史数据纳入RCT,特别是解决未测量的混因素带来的偏差.
- 通过更有效地利用历史数据,提高RCT因果推理的准确性和可靠性.
- 引入先前的倾向性评分集成自适应混合物 (PS-SAM) 作为适应性信息借贷的解决方案.
主要方法:
- 整合一个自适应混合物 (SAM) 之前与倾向性得分匹配和反向概率加权.
- 开发倾向性得分集成的SAM (PS-SAM) 先验,以减轻未测量的混因子带来的偏差.
- 使用模拟研究来评估PS-SAM之前的操作特性.
主要成果:
- PS-SAM先验证明了稳定性,在没有测量不到的混因素存在时,产生了公正的因果估计.
- 在未测量的混因子存在的情况下,PS-SAM先验提供了明显更少偏差的治疗效果估计和改进的I型错误控制.
- 模拟结果证实了PS-SAM在适应性信息借用之前的理想操作特性.
结论:
- 预先的PS-SAM方法提供了一个强大的方法来将历史数据整合到RCT中,有效地处理未测量的混.
- 这种方法通过允许适应性信息借用来改善因果推断,从而使得治疗效果估计更可靠.
- 拟议的方法可以通过R包"SAMprior"访问.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
14
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
14
Weighted Mean
4.8K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
4.8K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
The Availability Heuristic
5.8K
A heuristic is a general problem-solving framework (Tversky & Kahneman, 1974). You can think of these as mental shortcuts that are used to solve problems. Different types of heuristics are used in different types of situations, and the impulse to use a heuristic occurs when one of five conditions is met (Pratkanis, 1989):
5.8K
Variation
6.7K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.7K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K


