Evaluating Ensemble-Based Machine Learning Models for Diagnosing Pediatric Acute Appendicitis: Insights from a
Zeynep Kucukakcali1, Sami Akbulut1,2, Cemil Colak1
1Department of Biostatistics and Medical Informatics, Inonu University Faculty of Medicine, 44280 Malatya, Turkey.
Insights
Machine learning models accurately classify pediatric acute appendicitis (AAP) subtypes. Random Forest and XGBoost show promise in improving diagnosis and patient outcomes by distinguishing between negative, uncomplicated, and complicated cases.
Area of Science:
- Computational biology and bioinformatics
- Pediatric surgery and emergency medicine
- Artificial intelligence in healthcare
Background:
- Pediatric acute appendicitis (AAP) diagnosis is challenging, with misclassification risking delayed treatment or unnecessary surgery.
- Accurate classification into negative, uncomplicated, and complicated AAP is crucial for appropriate pediatric care.
- Existing diagnostic methods for AAP require enhancement to improve precision and patient outcomes.
Purpose of the Study:
- To evaluate and compare the diagnostic accuracy of five machine learning (ML) models for classifying pediatric AAP subtypes.
- To identify the most effective ML models for distinguishing between negative, uncomplicated, and complicated pediatric appendicitis.
- To assess the role of specific laboratory biomarkers in the ML-based classification of AAP.
Main Methods:
- Retrospective analysis of 590 pediatric patients diagnosed with AAP.
- Inclusion of demographic data and laboratory parameters (CRP, WBC, neutrophils, lymphocytes, appendiceal diameter) as features.
- Training and testing of five ensemble ML models (AdaBoost, XGBoost, Stochastic Gradient Boosting, Bagged CART, Random Forest) using cross-validation.
Main Results:
- Random Forest achieved 90.7% accuracy, 100% sensitivity, and 61.5% specificity for negative vs. uncomplicated AAP.
- XGBoost demonstrated superior performance for complicated AAP with 97.3% accuracy, 100% sensitivity, and 78.3% specificity.
- Neutrophil count, appendiceal diameter, and WBC levels were identified as the most influential predictive biomarkers.
Conclusions:
- Machine learning models, specifically Random Forest and XGBoost, show significant potential in aiding pediatric AAP diagnosis.
- ML-based decision support tools can enhance clinical judgment, leading to improved diagnostic accuracy and patient outcomes.
- Future research should focus on multi-center validation, integrating imaging data, and improving model interpretability for clinical adoption.
Abstract:
Background: Pediatric acute appendicitis (AAP) is a common cause of abdominal pain in children, yet accurate classification into negative, uncomplicated, and complicated forms remains clinically challenging. Misclassification may lead to unnecessary surgeries or delayed treatment. This study aims to evaluate and compare the diagnostic accuracy of five machine learning models (AdaBoost, XGBoost, Stochastic Gradient Boosting, Bagged CART, and Random Forest) for classifying pediatric AAP subtypes. Methods: In this retrospective observational study, a dataset of 590 pediatric patients was analyzed. Demographic information and laboratory parameters-including C-reactive protein (CRP), white blood cell (WBC) count, neutrophils, lymphocytes, and appendiceal diameter-were included as features. The cohort consisted of negative (19.8%), uncomplicated (49.2%), and complicated (31.0%) AAP cases. Five ensemble machine learning models (AdaBoost, XGBoost, Stochastic Gradient Boosting, Bagged CART, and Random Forest) were trained on 80% of the dataset and tested on the remaining 20%. Model performance was evaluated using accuracy, sensitivity, specificity, and F1 score, with cross-validation employed to ensure result stability. Results: Random Forest demonstrated the highest overall accuracy (90.7%), sensitivity (100.0%), and specificity (61.5%) for distinguishing negative and uncomplicated AAP cases. Meanwhile, XGBoost outperformed other models in identifying complicated AAP cases, achieving an accuracy of 97.3%, sensitivity of 100.0%, and specificity of 78.3%. The most influential biomarkers were neutrophil count, appendiceal diameter, and WBC levels, highlighting their predictive value in AAP classification. Conclusions: ML models, particularly Random Forest and XGBoost, exhibit strong potential in aiding pediatric AAP diagnosis. Their ability to accurately classify AAP subtypes suggests that ML-based decision support tools can complement clinical judgment, improving diagnostic precision and patient outcomes. Future research should focus on multi-center validation, integrating imaging data, and enhancing model interpretability for broader clinical adoption.
Related Concept Videos
Appendicitis-II: Diagnostic Studies and Management
Diagnosing Appendicitis
It requires a multifaceted approach, starting with a detailed physical examination to pinpoint the location and nature of the pain and identify any associated symptoms. Laboratory tests play a crucial role. A complete Blood Count (CBC) typically reveals leukocytosis (an increased number of...
Appendicitis-I: Introduction
Etiology: Appendicitis can arise from various causes, primarily rooted in the obstruction of the appendix lumen. Factors contributing to this obstruction include fecal accumulation, lymphoid hyperplasia and, in...


