Explainable AI based cervical cancer prediction using FSAE feature engineering and H2O AutoML
Panneerselvam Karthikeyan1, I Malaserene2, E Deepakraj1
1School of Computer Science Engineering and Information Systems, Vellore Institute of Technology, Vellore, India.
Scientific Reports
|November 18, 2025
Summary
This study introduces a hybrid machine learning (ML) framework for predicting cervical cancer risk. The model combines automated machine learning (AutoML) with feature engineering and explainable AI (XAI) for improved accuracy and interpretability.
Area of Science:
- Oncology
- Biomedical Informatics
- Machine Learning
Background:
- Cervical cancer, primarily caused by Human Papillomavirus (HPV), poses a significant global health challenge, necessitating accurate early prediction for improved patient outcomes.
- Existing machine learning (ML) and deep learning (DL) models for disease prediction often lack interpretability and require extensive datasets, while conventional diagnostic methods can be costly and complex.
- There is a critical need for advanced, interpretable, and efficient methods for cervical cancer risk prediction to support clinical decision-making.
Purpose of the Study:
- To develop and validate a hybrid ML framework for enhanced cervical cancer risk prediction.
- To integrate automated feature extraction and selection with AutoML for improved model performance.
- To incorporate explainable AI (XAI) techniques, including LIME and SHAP, to enhance model transparency and clinical trust.
Main Methods:
- A hybrid framework combining H2O AutoML with autoencoder-based feature extraction and Fisher Score-based feature selection was developed.
- Exploratory data analysis (EDA) and dimensionality reduction were performed using a stacked autoencoder.
- Local Interpretable Model-Agnostic Explanations (LIME) and SHAP were utilized for instance-level prediction interpretation.
Main Results:
- The selected deep learning model achieved high performance on the training dataset with 95.24% accuracy and an AUC of 98.10.
- Cross-validation demonstrated the model's robustness, with consistent AUC and log loss values.
- The model exhibited a low overall error rate of 4.14%, with specific error rates for actual negatives (5.75%) and actual positives (2.59%) at an F1 threshold of 0.517.
Conclusions:
- The proposed hybrid ML framework effectively enhances the predictive power and interpretability of cervical cancer risk models.
- Combining AutoML with advanced feature engineering and XAI offers a scalable and clinically relevant solution for decision support.
- The developed model provides actionable insights for clinicians, paving the way for more personalized and effective cervical cancer management.


