基于机器学习的预测模型的阶级失衡纠正的危害:一项模拟研究
Alex Carriero1, Kim Luijken1, Anne de Hond1
1Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht, The Netherlands.
Statistics in medicine
|January 27, 2025
概括
在机器学习模型中纠正类失衡可能会损害校准. 没有失衡校正的模型始终显示出更好或相同的校准,避免在临床预测中过度估计风险.
科学领域:
- 机器学习 机器学习
- 生物统计学 生物统计学
- 临床信息学 临床信息学
背景情况:
- 风险预测模型对于临床决策至关重要,需要准确的校准.
- 医疗保健数据经常表现出阶级不平衡,导致研究人员应用纠正.
- 这些失衡纠正对模型校准的影响仍然不清楚.
研究的目的:
- 调查阶级失衡纠正对机器学习模型校准的影响.
- 为了在各种场景中比较具有和没有失衡校正的模型之间的校准性能.
主要方法:
- 利用广泛的蒙特卡洛模拟来评估样本之外的预测性能.
- 在不同的数据生成条件下评估机器学习算法 (样本大小,预测器,事件分数).
- 通过使用MIMIC-III数据的案例研究来说明研究结果.
主要成果:
- 在没有进行类不平衡校正的情况下开发的模型始终表现出优越或同等的校准.
- 失衡的纠正导致了校准错误,其特点是过度估计风险.
- 再校准并不总是解决失衡纠正引入的校准错误.
结论:
- 对于临床预测模型来说,类失衡校正并不普遍要求.
- 应用失衡校正可能会损害模型校准和可靠性.
- 为个人风险估计优先考虑模型校准而不是自动纠正失衡.
相关概念视频
Survival Tree
55
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
55
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Mechanistic Models: Compartment Models in Individual and Population Analysis
26
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
26
Accuracy and Errors in Hypothesis Testing
171
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
171
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K


