相关实验视频
Updated: Jan 11, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
2.5K
预测性能指标的行为与罕见事件的行为
Emily Minus1, R Yates Coley2, Susan M Shortreed3
1Department of Biostatistics, University of Washington, Seattle, WA, USA.
Journal of clinical epidemiology
|November 12, 2025
概括
接收器运行特征曲线 (AUC) 下的面积对于罕见事件是可靠的,如果至少有1000个事件. 这一发现对于在大型医疗保健数据集中评估自杀风险预测模型至关重要.
科学领域:
- 生物统计学 生物统计学
- 医疗信息学 医疗信息学
- 临床预测建模临床预测建模
背景情况:
- 接收器运行特征曲线 (AUC) 下的面积是二进制结果预测模型的标准度量.
- 有关AUC在罕见事件环境中的可靠性存在担忧,这在临床实践中很常见.
- 自杀风险预测模型至关重要,但需要强有力的绩效评估.
研究的目的:
- 为了调查在罕见事件设置中AUC的不稳定性是由于事件数或事件速率.
- 评估其他预测性能指标的行为,如灵敏度,特异性,积极的预测值和准确性.
- 确定在罕见事件情景中可靠估计AUC所需的事件的最小数量.
主要方法:
- 使用基于真实世界健康记录数据的等离子模式方法进行的模拟研究.
- 研究数据集包括149个预测因素和一个罕见的结局 (自杀未遂),事件发生率为0.92%.
- 模拟改变了数据集大小,以分析不同事件计数下的AUC偏差和差异.
主要成果:
- 在罕见事件设置中的AUC不稳定性主要是由事件的总数驱动的,而不是事件率.
- 对AUC的近零偏差观察到大约有1000个事件.
- 灵敏度的表现取决于事件的数量,而特异性取决于非事件;PPV和准确性受事件率的影响.
结论:
- 当存在足够数量的事件 (例如1000个) 时,AUC是罕见事件预测模型的可靠性能指标,包括自杀风险模型.
- 大规模的医疗保健数据库可以支持对罕见事件的可靠AUC评估.
- 事件的数量,而不是事件率,是罕见事件预测中AUC稳定的关键因素.
相关概念视频
Unusual Results
3.7K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
3.7K
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K
Random Error
7.8K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
7.8K
Probability in Statistics
22.0K
Probability is the likelihood of an event occurring. The term event is defined as a collection of results of a procedure. An event is a simple event when an outcome cannot be divided into simpler parts.
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
22.0K
Expected Frequencies in Goodness-of-Fit Tests
7.1K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
7.1K

