XGBoost-SHAP-based interpretable diagnostic framework for knee osteoarthritis: a population-based retrospective
Zijuan Fan1,2, Wenzhu Song3, Yan Ke4
1Department of Orthopaedic Surgery, The First Affiliated Hospital, Zhejiang University School of Medicine, Qingchun Road No. 79, Hangzhou, China.
Arthritis Research & Therapy
|December 19, 2024
Summary
Machine learning models can diagnose knee osteoarthritis (KOA) using routine data. Joint pain experience emerged as the most significant factor for KOA diagnosis.
Area of Science:
- Orthopedics
- Medical Informatics
- Machine Learning
Background:
- Knee osteoarthritis (KOA) is a prevalent degenerative joint disease.
- Accurate and early diagnosis of KOA is crucial for effective management and prevention strategies.
Purpose of the Study:
- To develop an interpretable machine learning (ML) model for diagnosing KOA using routine demographic and clinical data.
- To identify key features contributing to KOA diagnosis.
Main Methods:
- A retrospective, population-based cohort study using questionnaire data from the Wu Chuan KOA Study.
- Feature selection, class balancing, and comparison of four ML classifiers (XGBoost with Boruta identified as best).
- Model performance evaluated using AUC, G-means, and F1 scores; feature importance determined by Shapley values.
Main Results:
- The study included 1188 participants, with 26.3% diagnosed with KOA.
- XGBoost with Boruta achieved the highest performance (AUC: 0.758, G-means: 0.800, F1: 0.703).
- The average experience of joint pain was identified as the most important feature for KOA diagnosis among the top 17 ranked features.
Conclusions:
- Machine learning models are effective in identifying crucial factors for KOA diagnosis.
- The findings can inform new prevention strategies for knee osteoarthritis.
- Further validation of this ML approach is recommended.
Related Concept Videos
Bias in Epidemiological Studies
153
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
153
Statistical Methods for Analyzing Epidemiological Data
299
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
299
Statistical Software for Data Analysis and Clinical Trials
492
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
492
Confounding in Epidemiological Studies
143
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
143
Genome-wide Association Studies-GWAS
12.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.4K
Biostatistics: Overview
220
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
220


