An Interpretable Machine Learning Framework for Rare Disease: A Case Study to Stratify Infection Risk in Pediatric
Irfan Al-Hussaini1,2, Brandon White1,3, Armon Varmeziar1,3
1Laboratory for Pathology Dynamics, Georgia Institute of Technology and Emory University, Atlanta, GA 30332, USA.
Abstract:
Background: Datasets on rare diseases, like pediatric acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL), have small sample sizes that hinder machine learning (ML). The objective was to develop an interpretable ML framework to elucidate actionable insights from small tabular rare disease datasets. Methods: The comprehensive framework employed optimized data imputation and sampling, supervised and unsupervised learning, and literature-based discovery (LBD). The framework was deployed to assess treatment-related infection in pediatric AML and ALL. Results: An interpretable decision tree classified the risk of infection as either "high risk" or "low risk" in pediatric ALL (n = 580) and AML (n = 132) with accuracy of ∼79%. Interpretable regression models predicted the discrete number of developed infections with a mean absolute error (MAE) of 2.26 for bacterial infections and an MAE of 1.29 for viral infections. Features that best explained the development of infection were the chemotherapy regimen, cancer cells in the central nervous system at initial diagnosis, chemotherapy course, leukemia type, Down syndrome, race, and National Cancer Institute risk classification. Finally, SemNet 2.0, an open-source LBD software that links relationships from 33+ million PubMed articles, identified additional features for the prediction of infection, like glucose, iron, neutropenia-reducing growth factors, and systemic lupus erythematosus (SLE). Conclusions: The developed ML framework enabled state-of-the-art, interpretable predictions using rare disease tabular datasets. ML model performance baselines were successfully produced to predict infection in pediatric AML and ALL.
Insights
This study developed an interpretable machine learning (ML) framework to predict treatment-related infections in pediatric acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL), achieving accurate risk classification and identifying key predictive features.
Area of Science:
- Computational biology
- Machine learning in rare diseases
- Pediatric oncology
Background:
- Rare disease datasets, such as those for pediatric acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL), present challenges for machine learning (ML) due to small sample sizes.
- Developing effective ML models for rare diseases requires specialized frameworks that can handle limited data.
Purpose of the Study:
- To create an interpretable ML framework for extracting actionable insights from small, tabular rare disease datasets.
- To apply this framework to predict treatment-related infections in pediatric AML and ALL.
Main Methods:
- The framework integrated optimized data imputation and sampling with supervised and unsupervised learning techniques.
- Literature-based discovery (LBD) using SemNet 2.0 was employed to identify additional predictive features.
- Interpretable ML models, including decision trees and regression models, were developed.
Main Results:
- An interpretable decision tree achieved ~79% accuracy in classifying infection risk (high/low) in pediatric ALL and AML.
- Regression models predicted the number of bacterial and viral infections with mean absolute errors of 2.26 and 1.29, respectively.
- Key predictive features included chemotherapy regimen, central nervous system involvement, leukemia type, and Down syndrome. LBD identified glucose, iron, and growth factors as additional predictors.
Conclusions:
- The developed ML framework provides state-of-the-art, interpretable predictions from rare disease datasets.
- This study establishes baseline ML model performance for predicting infections in pediatric AML and ALL.
- The framework demonstrates the potential for advancing ML applications in rare pediatric cancers.


