连续共变量和复杂回归模型的分类 - - 在交叉性研究中的朋友或敌人
Adrian Richter1, Sabina Ulbricht1, Sarah Brockhaus2
1Department of Prevention Research and Social Medicine, Institute for Community Medicine, University Medicine Greifswald, Greifswald, Germany.
Journal of clinical epidemiology
|April 24, 2024
概括
在交叉性研究中对连续变量进行分类会导致错误的发现. 使用贝叶斯信息标准 (BE-BIC) 和splines方法的倒向变量消除更强大,更可概括,更易于解释.
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
- 健康 公平 研究 健康 公平 研究
背景情况:
- 识别健康不平等需要了解个体特征的交叉点.
- 交叉性分析的现有方法包括使用一级和二级效应或分层的建模,这两种方法都对连续的共同变量进行了分类.
- 这些方法在识别真交叉点和避免假阳性的性能需要评估.
研究的目的:
- 为了比较两个新的交叉性分析方法与标准方法的性能.
- 基于识别真交点,假阳性率和对独立数据的概括性来评估方法.
- 确定将连续共变量分类对交叉性研究的影响.
主要方法:
- 使用R软件进行了模拟研究.
- 模拟了共变量 (年龄,性别,BMI,教育,糖尿病),以评估与骨质疏松症持续脆弱性得分的关联.
- 我们比较了三种方法: 1) 一级和二级效应的模型, 2) 分层, 3) 以贝叶斯信息标准 (BE-BIC) 和线条的向后变量消除.
- 模型在不同的样本大小和信号噪声比下进行了评估,使用引导重新抽样进行验证和平均平方误差 (MSE) 进行比较.
主要成果:
- 方法1和2在没有真正的共变量关联的模拟中产生了90%以上的虚假效应.
- 方法3 (BE-BIC) 在选择正确模型方面显示出更高的准确性 (36.5%至89.8%),以及更少的虚假效应,特别是在更大的样本大小的情况下.
- 与方法1和2相比,方法3在独立数据中显示MSE较低,并且在所有设置中产生较少的虚假结果.
结论:
- 将连续的共变量归类对交叉性研究是有害的,增加了虚假效应并降低了可解释性.
- BE-BIC方法 (方法3) 对虚假发现更具稳定性,并为独立数据提供更好的概括性.
- 对于可靠的交叉性研究,专注于描述相关的交叉差异,避免不可重复的发现至关重要.
相关概念视频
Multicompartment Models: Overview
138
Multicompartment models are mathematical constructs that depict how drugs are distributed and eliminated within the body. They segment the body into several compartments, symbolizing various physiological or anatomical areas connected through drug transfer processes such as absorption, metabolism, distribution, and elimination.
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
138
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Mechanistic Models: Compartment Models in Individual and Population Analysis
38
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
38
Strategies for Assessing and Addressing Confounding
94
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
94
Confounding in Epidemiological Studies
165
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
165
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K


