构建一个更好的模型:放弃厨房水槽回归
Stefan Kuhle1,2, Mary Margaret Brown3, Sanja Stanojevic4
1Institute for Medical Biostatistics, Epidemiology and Informatics (IMBEI), Johannes Gutenberg University Mainz, Mainz, Germany stefan.kuhle@uni-mainz.de.
Archives of disease in childhood. Fetal and neonatal edition
|December 10, 2023
概括
厨房水槽回归,使用p值来选择变量,对于医学研究来说是不可靠的. 定向非循环图 (DAG) 为构建更好的回归模型提供了一个强大的替代方案.
科学领域:
- 生物统计学 生物统计学
- 流行病学 流行病学
- 医学研究方法学 医学研究方法学
背景情况:
- 多变量回归模型在医学研究中对于了解疾病风险因素至关重要.
- 变量选择是回归建模中的关键步骤,影响结果的有效性.
- 当前的实践,如"厨房水槽回归",往往依赖于统计标准,而不是因果推理.
研究的目的:
- 批判性地评估"厨房水槽回归"方法用于变量选择.
- 确定使用p值或模型构建信息标准的固有陷和局限性.
- 为医学研究中的回归模型开发提出替代性,更强大的方法.
主要方法:
- 对"厨房水槽回归"方法进行批判性检查.
- 使用围产/新生儿医学的例子说明陷.
- 引入和应用定向环形图 (DAG) 用于因果分析.
主要成果:
- "厨房水槽回归"忽视了变量方向性,缺乏因果解释,膨胀了I型错误率,风险过度,并忽视了内容专业知识.
- 用"厨房水槽回归"方法确定了五个关键问题.
- DAG为理解和可视化因果关系提供了一个框架.
结论:
- "厨房水槽回归"是医学研究中选择变量的有缺陷的方法.
- 建议使用定向非循环图 (DAG) 来指导可变选择,以检查风险因子-结果关联.
- 对回归建模的更深思熟虑和更明智的方法,包括因果推理是必不可少的.
相关概念视频
Improving Translational Accuracy
11.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
515
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
515


