Identifying Data-Driven Clinical Subgroups for Cervical Cancer Prevention With Machine Learning: Population-Based,
Zhen Lu1, Binhua Dong2,3, Hongning Cai4
1School of Public Health (Shenzhen), Sun Yat-sen University, Shenzhen, China.
JMIR Public Health and Surveillance
|March 19, 2025
Summary
Machine learning identified 5 cervical cancer prevention subgroups with distinct precancer risks. This enables personalized strategies, prioritizing high-risk groups for colposcopy and scaling HPV screening for lower-risk groups.
Area of Science:
- Oncology
- Genomics
- Public Health
Background:
- Cervical cancer prevention (CCP) remains a significant global health challenge.
- Personalized, data-driven CCP strategies are needed to improve outcomes.
- Tailoring prevention to phenotypic profiles can reduce disease burden.
Purpose of the Study:
- Identify distinct cervical precancer and cancer risk subgroups using machine learning.
- Validate subgroup predictions across independent datasets.
- Propose a computational phenomapping strategy to enhance global CCP efforts.
Main Methods:
- Applied unsupervised machine learning to a deeply phenotyped cohort to identify CCP subgroups.
- Used weighted logistic regression to determine risks of cervical intraepithelial neoplasia (CIN2+ and CIN3+).
- Trained a supervised model for individual classification and validated it on an external cohort.
Main Results:
- Identified 5 distinct CCP subgroups from over 550,000 women.
- Subgroups CCP2-4 exhibited significantly higher risks for CIN2+ and CIN3+ compared to CCP1.
- Validated a triple strategy prioritizing high-risk subgroups (CCP3-4) for colposcopy and scaling HPV screening for CCP1-2.
Conclusions:
- Machine learning and electronic health records can enhance CCP strategies.
- Identifying key determinants of CIN2+/CIN3+ risk and classifying subgroups provides a data-driven foundation for tailored prevention.
- The proposed triple strategy offers a scalable tool to complement existing cervical cancer screening guidelines.


