可视化和量化新预测因子的有用性,按结果类分层:U-smile方法
Katarzyna B Kubiak1, Barbara Więckowska1, Elżbieta Jodłowska-Siewert2
1Department of Computer Science and Statistics, Poznan University of Medical Sciences, Poznan, Poland.
PloS one
|May 20, 2024
概括
新的U-smile方法可视评估对二进制分类的预测改进. 它有效地突出了预测器的有用性,在模型比较中表现优于ROC曲线.
科学领域:
- 统计 统计 统计 统计
- 机器学习 机器学习
- 生物统计学 生物统计学
背景情况:
- 二进制分类和预测在许多领域都至关重要.
- 评估预测改进现有的方法可能是有限的.
- 需要直观和定量工具来评估模型性能.
研究的目的:
- 引入和验证U-smile方法,一种新的图形和定量方法.
- 通过二进制结果类分层评估预测改进.
- 为了比较U-smile方法与接收器操作特征 (ROC) 曲线.
主要方法:
- 开发了U-smile方法,使用了新的系数和类似微笑的情节.
- 应用概率比测试来评估预测变化的意义.
- 在心脏病数据集和随机变量上使用后勤回归模型进行验证.
主要成果:
- 微笑方法始终为信息预测器生成微笑形图形.
- 在比较预测效应时,U微笑图表在ROC曲线上表现出更高的有效性.
- 该方法清楚地区分了事件和非事件的模型性能.
结论:
- U-smile方法提供了一个直观的预测值的视觉评估.
- 它有助于选择对二进制预测模型最有影响力的预测因素.
- 微笑法在预测评估之外有潜在的应用.
相关概念视频
Survival Tree
80
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
80
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Sensitivity, Specificity, and Predicted Value
289
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
289
Response Surface Methodology
120
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
120
Variation
6.8K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.8K
Review and Preview
7.3K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
7.3K


