An integrated and interpretable machine learning framework for Kawasaki disease diagnosis and risk prediction
Dandan Wang1, Fei Li2, Tingting Xie1
1Department of Pediatrics, The First Affiliated Hospital of University of Science and Technology of China, Division of Life Sciences and Medicine, University of Science and Technology of China, Hefei, China.
Background:
Early identification of Kawasaki disease (KD) and accurate prediction of its associated complications are critical for optimizing treatment strategies and improving clinical outcomes. While machine learning has shown promise in KD-related studies, most existing models are limited to single tasks with varying feature sets, lacking an integrated framework. This fragmentation hinders clinical applicability and constrains generalizability. Therefore, this study aimed to develop a unified and interpretable machine learning framework for KD diagnosis and risk prediction, with the goal of enhancing clinical relevance and real-world applicability.
Methods:
We retrospectively collected data from 2,133 febrile pediatric patients treated at The First Affiliated Hospital of University of Science and Technology of China between January 1, 2018, and December 31, 2022. After excluding patients older than 5 years or with incomplete records, a total of 919 cases-including both typical and atypical KD-were included. Using 29 common clinical features, we developed a unified light gradient boosting machine (LightGBM)-based model for KD diagnosis, intravenous immunoglobulin (IVIG) resistance prediction, and coronary artery lesion (CAL) risk assessment. The dataset was split into training and validation sets at an 8:2 ratio, and five-fold cross-validation was performed to ensure robustness. Model performance was evaluated using accuracy, area under the receiver operating characteristic (ROC) curve (AUC), sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Feature importance and model interpretability were assessed using SHapley Additive exPlanations (SHAP). To further assess its clinical utility, we compared the model's diagnostic performance with that of pediatric clinicians and an advanced large language model (ChatGPT).
Results:
The KD diagnostic task achieved an AUC of 0.999, with a sensitivity of 0.984 and specificity of 0.974. The IVIG resistance prediction task yielded an AUC of 0.888, sensitivity of 0.600, and specificity of 0.979. For CALs risk prediction, the AUC was 0.783, with a sensitivity of 0.529 and specificity of 0.984. SHAP analysis identified distinct sets of top-ranking features for each task, reflecting the underlying clinical heterogeneity. Notably, variables such as inflammatory markers, immune-related indicators, and characteristic clinical signs of KD consistently contributed to model predictions. In a comparative study, our model achieved accuracies of 0.900, 0.800, and 0.757 for KD diagnosis, IVIG resistance, and CAL prediction, respectively, consistently outperforming pediatricians with over 5 years of experience and ChatGPT, highlighting its potential as a clinical decision support tool.
Conclusions:
This study presents a unified machine learning framework that accurately supports KD diagnosis, IVIG resistance prediction, and CAL risk assessment. By leveraging a common set of clinical features, the model enhances clinical applicability and lays the groundwork for task-specific management and precision intervention in KD.
Related Concept Videos
Statistical Software for Data Analysis and Clinical Trials
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...


