膝关节骨关节炎的基于XGBoost-SHAP的可解释诊断框架:基于人口的回顾性队列研究
Zijuan Fan1,2, Wenzhu Song3, Yan Ke4
1Department of Orthopaedic Surgery, The First Affiliated Hospital, Zhejiang University School of Medicine, Qingchun Road No. 79, Hangzhou, China.
Arthritis research & therapy
|December 19, 2024
概括
机器学习模型可以使用常规数据诊断膝关节关节炎 (KOA). 关节疼痛经历成为KOA诊断中最重要的因素.
科学领域:
- 整形外科 整形外科 整形外科
- 医疗信息学 医疗信息学
- 机器学习 机器学习
背景情况:
- 膝关节骨关节炎 (KOA) 是一种普遍存在的退行性关节疾病.
- 准确和早期诊断KOA对于有效的管理和预防策略至关重要.
研究的目的:
- 开发一种可解释的机器学习 (ML) 模型,使用常规的人口和临床数据来诊断KOA.
- 为了确定有助于KOA诊断的关键特征.
主要方法:
- 一个回顾性,以人口为基础的队列研究,使用来自Wu Chuan KOA研究的问卷数据.
- 特性选择,类平衡和四个ML分类器的比较 (XGBoost与Boruta被确定为最好的).
- 模型性能使用AUC,G-平均值和F1分数进行评估;特征重要性由Shapley值确定.
主要成果:
- 该研究包括1188名参与者,其中26.3%被诊断为KOA.
- 使用Boruta的XGBoost获得了最高的性能 (AUC:0.758,G-平均值:0.800,F1:0.703).
- 关节疼痛的平均体验被确定为KOA诊断中最重要的特征,在排名前17个特征中.
结论:
- 机器学习模型有效地识别了KOA诊断的关键因素.
- 这些发现可以为膝关节骨关节炎的新预防策略提供信息.
- 建议进一步验证这一ML方法.
相关概念视频
Bias in Epidemiological Studies
153
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
153
Statistical Methods for Analyzing Epidemiological Data
299
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
299
Statistical Software for Data Analysis and Clinical Trials
492
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
492
Confounding in Epidemiological Studies
143
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
143
Genome-wide Association Studies-GWAS
12.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.4K
Biostatistics: Overview
220
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
220


