Related Experiment Video
Updated: Dec 19, 2025

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
Mining incomplete clinical data for the early assessment of Kawasaki disease based on feature clustering and
Haolin Wang1, Xuhai Tan2, Zhilin Huang3
1College of Medical Informatics, Chongqing Medical University, Chongqing, 400016, China.
Insights
Early diagnosis of Kawasaki disease (KD) is challenging. A new data-driven method using convolutional neural networks (CNNs) achieved 97% accuracy in identifying KD from electronic health records, improving early detection.
Area of Science:
- Pediatric Cardiology
- Medical Informatics
- Artificial Intelligence in Medicine
Background:
- Kawasaki disease (KD) is a primary cause of acquired heart disease in children.
- Early diagnosis and treatment of KD are crucial to prevent severe cardiac complications like coronary aneurysms.
- Current diagnostic challenges stem from unknown pathogenesis and the absence of specific diagnostic markers.
Purpose of the Study:
- To develop and evaluate a data-driven approach for the early assessment of Kawasaki disease using electronic health records.
- To address challenges posed by incomplete clinical data and group-based missing patterns.
- To leverage advanced machine learning techniques for improved diagnostic accuracy.
Main Methods:
- Utilized a cohort of 10,367 patients from electronic health records.
- Developed a novel method integrating feature clustering for matrix-based representation.
- Employed convolutional neural networks (CNNs) for feature extraction and fusion, exploiting multi-source data structure.
- Integrated missing data imputation techniques.
Main Results:
- The proposed method achieved a superior Area Under the Curve (AUC) of 0.97.
- Demonstrated significantly higher accuracy compared to benchmark methods in early KD assessment.
- Successfully addressed issues of incomplete clinical data and complex missing data patterns.
Conclusions:
- The developed data-driven method shows significant potential for improving clinical data mining in pediatric cardiology.
- Matrix-based feature representation and CNN-based feature extraction are effective for handling incomplete clinical data.
- This approach can support medical decision-making for earlier and more accurate Kawasaki disease diagnosis.
Abstract:
Kawasaki disease (KD) is the leading cause of acquired heart disease in children. Its prompt treatment can effectively lower the risk of severe complications, such as coronary aneurysms. However, accurately diagnosing KD at its early stage is impracticable given its unknown pathogenesis and lack of pathognomonic features. In this study, we investigated data-driven approaches by using a cohort of 10,367 patients extracted from electronic health records for early KD assessment. The incompleteness of clinical data presents group-based missing patterns associated with different clinical assessment measures. To address this problem, we developed a method integrating feature clustering to enable matrix-based representation and convolutional neural networks (CNN) for feature extraction and fusion to explicitly exploit the multi-source data structure. Integrating missing data imputation methods with the proposed method demonstrated superior accuracy (an AUC of 0.97) compared with a number of benchmark methods. The present method shows potential to improve clinical data mining. Our study highlighted the feasible utilization of matrix-based feature representation and CNN-based feature extraction for incomplete clinical data mining to support medical decision-making.