Performance of Machine Learning Models for Prognosis Prediction in Oral Cavity Squamous Cell Carcinoma: A Systematic
Sammy Y Gao1, Jonathan M Hughes1, Shaun A Nguyen1
1Department of Otolaryngology-Head and Neck Surgery, Medical University of South Carolina, Charleston, SC 29425, USA.
Abstract:
Background/Objectives: Machine learning (ML) models have increasingly been applied to prognostic prediction in oral cavity squamous cell carcinoma (OCSCC), though their performance and methodological quality remain variably reported. This systematic review evaluated contemporary ML-based prognostic models for clinically relevant OCSCC outcomes, with emphasis on independently validated studies. Methods: A systematic review was conducted according to PRISMA guidelines using PubMed, Scopus, Cochrane Library, and CINAHL databases through 1 December 2025. Studies evaluating ML or artificial intelligence prognostic models in adult OCSCC patients were included. Outcomes included overall survival, recurrence, disease-free survival, recurrence-free survival, disease-specific survival, cancer-specific survival, progression, and nodal metastasis. Data extraction and risk-of-bias assessment using PROBAST + AI were performed independently by reviewers. Results: Forty studies comprising 105,619 patients met inclusion criteria. ML architectures included random forests, support vector machines, gradient boosting methods, neural networks, and deep learning frameworks. Most models incorporated clinical and pathologic variables, while many integrated radiologic, immunologic, or genomic features. For overall survival prediction, independently validated models generally demonstrated AUCs between 0.80 and 0.90. Recurrence prediction models similarly showed favorable discrimination, with most externally validated studies reporting acceptable predictive performance. Additional prognostic endpoints including disease-free survival, progression, and nodal metastasis demonstrated AUCs ranging from 0.70 to 0.90. Common methodological limitations included retrospective design, small sample size, inadequate external validation, and risk of overfitting. Conclusions: ML-based prognostic models in OCSCC demonstrate generally favorable predictive performance across survival and recurrence outcomes. However, substantial heterogeneity in methodology and limited external validation continue to restrict clinical implementation. Future work should prioritize prospective multicenter validation, standardized reporting, and reproducible modeling frameworks.
