在二进制回归模型中确定抵消值以调整错误分类错误
1Department of Public Health Sciences, 532 Edwards Hall, Clemson University, Clemson, SC 29634, USA.
International journal of environmental research and public health
|February 27, 2026
概括
本研究引入了一种方法来纠正在临床研究中使用的代理措施中的错误分类错误引起的偏见统计推断. 该方法使用回归模型中的偏移值来确保准确的效应估计,提高数据可靠性.
科学领域:
- 生物统计学 生物统计学
- 临床流行病学临床流行病学
- 医疗保健服务研究 医疗服务研究
背景情况:
- 在研究中收集黄金标准的结果措施往往是昂贵的,并且在后勤上具有挑战性.
- 替代或代理措施经常被使用,但可能导致错误分类错误和偏见的统计推断.
- 现有的方法很难完全解释不完美的结果措施引入的偏差.
研究的目的:
- 开发一种统计方法,以纠正因代理措施中的错误分类错误而导致的影响估计的偏差.
- 为了确定一般化二进制回归模型的适当的偏移值,以消除偏差.
- 通过模拟研究验证拟议的方法.
主要方法:
- 利用了包含偏移值的通用二进制回归模型.
- 基于来自验证样本 (内部或外部) 的错误分类错误率计算的偏移值.
- 采用模拟研究来证明和验证各种效果测量的偏差校正方法.
主要成果:
- 提出的方法成功确定了偏移值,以消除风险差异,相对风险和赔率比率估计中的偏差.
- 模拟研究证实了对偏移调整模型的准确性.
- 调整后的模型中的点估计和标准误差都被证明是无偏见的.
结论:
- 偏移调整回归建模方法有效地纠正了代理措施中的错误分类错误.
- 这种方法在研究中提高了统计推断的可靠性,在研究中,金标准措施是不可行的.
- 这些发现对提高使用替代结果数据的研究有效性有重大影响.
相关概念视频
Regression Toward the Mean
7.2K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.2K
Variation
8.2K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
8.2K
Regression Analysis
8.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.7K
Residuals and Least-Squares Property
9.7K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.7K
Multiple Regression
4.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.2K
Strategies for Assessing and Addressing Confounding
489
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
489


