Explainable Machine Learning for Head and Neck Cancer Risk Stratification
Amr Sayed Ghanem1, Róbert Bata1, Marianna Móré2
1Department of Epidemiology, Faculty of Health Sciences, University of Debrecen, 4028 Debrecen, Hungary.
Cancers
|July 28, 2026
Summary
Explainable machine learning models using electronic health records can accurately identify patients at high risk for head and neck cancers. This approach aids in early detection within routine clinical care.
Area of Science:
- Oncology
- Medical Informatics
- Machine Learning
Background:
- Head and neck cancers often present at advanced stages, leading to poor prognoses despite treatment advancements.
- Early cancer detection is crucial for improving patient outcomes but remains a challenge in routine clinical practice.
- Electronic health records (EHRs) offer a rich source of longitudinal data for identifying clinical signals, but machine learning (ML) model interpretability and integration are barriers.
Purpose of the Study:
- To develop and compare explainable ML models for early risk stratification of head and neck cancers using routine EHR data.
- To assess the predictive performance and clinical interpretability of different survival modeling approaches.
- To evaluate the potential utility of these models for opportunistic patient identification within existing healthcare workflows.
Main Methods:
- A retrospective cohort study of 157,031 patients treated between 2007 and 2022, with head and neck cancer identified via ICD-10 codes C01-C14.
- Feature reduction from 1397 to 91 variables using variance filtering and elastic net penalized Cox regression.
- Development and comparison of three survival models: CoxNet, Random Survival Forest, and XGBoost with a Cox objective.
Main Results:
- The XGBoost model achieved the highest predictive performance (concordance index: 0.916), outperforming Random Survival Forest (0.892) and CoxNet (0.886).
- All models demonstrated acceptable calibration, with clear risk stratification into low, medium, and high-risk groups.
- SHapley Additive exPlanations (SHAP) revealed that predictions were driven by demographic factors, laboratory markers, and diagnosis codes, reflecting clinical relevance.
Conclusions:
- Explainable ML models applied to routine EHR data can provide accurate and interpretable risk stratification for head and neck cancers.
- These models hold potential for the opportunistic early identification of high-risk individuals within current healthcare pathways.
- The findings support the integration of explainable AI into clinical workflows to enhance cancer risk assessment and patient management.

