利用沙普利的附加解释,在糖尿病预测的整体模型中选择特征
Prasant Kumar Mohanty1, Sharmila Anand John Francis2, Rabindra Kumar Barik3
1Department of Computer Science and Engineering, National Institute of Technology, Meghalaya 793003, India.
Bioengineering (Basel, Switzerland)
|January 8, 2025
概括
这项研究使用Shapley添加式解释 (SHAP) 与机器学习来识别关键糖尿病风险因素. 专注于前三项功能显著提高了临床应用的预测准确性和效率.
科学领域:
- 医疗信息学 医疗信息学
- 计算生物学 计算生物学
- 公共卫生 公共卫生
背景情况:
- 糖尿病是一个全球性的健康挑战,在印度加剧了与城市化有关的生活方式变化.
- 有效的糖尿病管理需要先进的预防策略和技术整合.
研究的目的:
- 提高糖尿病预测模型的准确性和效率.
- 用SHAP值识别和验证使用糖尿病风险最有影响力的特征.
- 评估特征选择对模型性能的影响.
主要方法:
- 将沙普利增量解释 (SHAP) 与整体机器学习模型集成.
- 使用三个特征集评估模型性能:所有特征,前三大影响力特征,以及不包括前三大特征的特征.
- 对十种不同的机器学习模型进行比较分析.
主要成果:
- 使用前三大最具影响力的特征的模型表现出卓越的性能.
- 集体模型在专注于顶级功能时,在大多数指标上取得了更好的表现.
- 排除前三大特征导致预测性能的显著下降.
结论:
- 使用SHAP值的有针对性的特征选择对于改善糖尿病预测模型是有效的.
- 这种方法提高了临床应用的模型效率和稳定性.
- 识别关键的预测特征对于准确和资源高效的糖尿病管理至关重要.
相关概念视频
Sensitivity, Specificity, and Predicted Value
180
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
180
Diabetes Mellitus: Type 2 and Gestational
2.2K
Type 2 diabetes, characterized by insulin resistance, arises when the insulin receptors on cells lose responsiveness to insulin, diminishing the cell's capacity to take up glucose, resulting in elevated blood glucose levels. To receive a diagnosis of Type 2 diabetes, a series of blood glucose tests are necessary to assess whether the blood glucose falls within normal parameters. If the result is out of the normal range, a patient may be diagnosed as prediabetic or diabetic, depending on the...
2.2K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Diabetes Mellitus: Overview and Type I Subtype
2.4K
Diabetes mellitus is a chronic metabolic disorder characterized by high blood glucose levels due to inadequate insulin production, insulin resistance, or both. The condition affects millions worldwide and can significantly impact their health and quality of life.
Type 1 diabetes is an autoimmune disease in which the immune system mistakenly attacks and destroys the insulin-producing beta cells in the pancreas. As a result, the body is unable to produce sufficient insulin, and individuals with...
Type 1 diabetes is an autoimmune disease in which the immune system mistakenly attacks and destroys the insulin-producing beta cells in the pancreas. As a result, the body is unable to produce sufficient insulin, and individuals with...
2.4K
Cancer Survival Analysis
328
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
328
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K


