Related Experiment Video
Updated: Jan 9, 2026

Experimental Model to Evaluate Resolution of Pneumonia
Published on: February 17, 2023
Development and validation of an interpretable machine learning model using routine laboratory biomarkers to stratify
Wei Cui1, Xinlv Zhang2, Yang Chen2
1Anhui Provincial Children's Hospital, Hefei, Anhui, China; National Children's Regional Medical Center, Hefei, Anhui, China; Anhui Clinical Medical Research Center for Child Health and Diseases, Hefei, Anhui, China; Anhui Institute of Pediatric Medicine, Hefei, Anhui, China.
Insights
This study developed an interpretable machine learning model to diagnose severe pneumonia in children and predict their risk of progression using routine lab tests. The CatBoost model aids early intervention, especially in resource-limited settings.
Area of Science:
- Pediatric critical care medicine
- Biomarker discovery
- Machine learning in healthcare
Background:
- Severe pneumonia is a major global cause of mortality in children under five.
- Accurate risk stratification tools for early identification of severe pneumonia are lacking.
Purpose of the Study:
- To develop an interpretable machine learning (IML) model for diagnosing severe pneumonia at admission.
- To predict the risk of pneumonia progression during hospitalization using routine laboratory biomarkers.
Main Methods:
- Retrospective analysis of 85,886 children with pneumonia from a Chinese tertiary hospital (2013-2023).
- Utilized 57 laboratory parameters from electronic health records for model development.
- Evaluated nine machine learning algorithms, focusing on CatBoost, with SHapley Additive exPlanations (SHAP) for interpretability.
Main Results:
- The CatBoost model, using 11 laboratory features, achieved an AUC of 0.879 for diagnosis and 0.839 for progression prediction.
- Optimized key feature thresholds (e.g., chloride ≤ 99 mmol/L) using Youden's index.
- A real-time web application with case-level interpretability was developed.
Conclusions:
- An interpretable CatBoost model effectively stratifies pediatric severe pneumonia risk using routine laboratory data.
- Clinical implementation via a web tool can support early intervention, particularly in resource-limited settings.
- External validation is recommended for broader applicability.
Introduction:
Severe pneumonia is a leading infectious cause of mortality in children under 5 years globally. Early identification of high-risk cases remains challenging due to the lack of reliable stratification tools.
Objectives:
This study aimed to develop an interpretable machine learning (IML) model using routine laboratory biomarkers for simultaneous diagnosis of severe pneumonia at admission and prediction of progression risk during hospitalization.
Methods:
This retrospective cohort study analyzed 85,886 children with pneumonia from a Chinese tertiary hospital (2013-2023). Two matched cohorts were established: Cohort I (n = 7,132) for admission diagnosis, and Cohort II (n = 1,064) for progression prediction. Fifty-seven laboratory parameters collected within 24 h of admission were extracted from electronic health records. Nine machine learning (ML) algorithms underwent systematic evaluation. Model performance was assessed using the area under the receiver-operating-characteristic curve (AUC), among other metrics. Model interpretability was achieved via SHapley Additive exPlanations (SHAP) analysis.
Results:
We evaluated the performance of nine machine learning algorithms and the CatBoost model incorporating 11 laboratory features demonstrated superior performance (AUC: 0.879 for admission diagnosis; 0.839 for progression prediction). Key feature thresholds were optimized using Youden's index (e.g., chloride ≤ 99 mmol/L). A real-time web application with case-level interpretability was deployed.
Conclusion:
This interpretable CatBoost model accurately stratifies pediatric severe pneumonia risk using routine laboratory data. Clinical implementation via the web tool may facilitate early intervention in resource-limited settings, though extensive external validation is warranted.
