Related Experiment Videos
Predicting high-grade lesions and cervical cancer with supervised machine learning algorithms
Laura C López-Alza1, Dauris L Mejía-Pérez1, Jairo Amaya-Guio1
1Universidad Nacional de Colombia, Department of Obstetrics and Gynecology, Sede Bogotá, Colombia.
Objective:
To evaluate machine learning models for predicting high-grade squamous intra-epithelial lesions and cervical intra-epithelial neoplasia grade 2 or higher using sociodemographic and clinical parameters.
Methods:
A retrospective cohort study was conducted using data from the Subred Norte (Bogotá, Colombia), a public health network, between 2018 and 2024. Records lacking sociodemographic data were excluded. To address the 10% class imbalance, random under-sampling was applied to train balanced models, which were then compared with the original imbalanced models.
Results:
The final sample included 4367 patients (aged 21 to 65 years) with cytology, colposcopy, or biopsy reports. The real-world prevalence of histopathologically confirmed cervical intra-epithelial neoplasia grade 2 or higher in this cohort was 10%. Five machine learning classification methods were tested across 2 datasets: 2018-2022 (n = 3139, excluding human papillomavirus history due to under-reporting) and 2023-2024 (n = 1228, including human papillomavirus history). On the original imbalanced dataset (without human papillomavirus history), the random forest model achieved an area under the receiver operating characteristic curve of 0.91 and a specificity of 99%, but had a sensitivity of only 41%. After applying random under-sampling, the random forest model's sensitivity improved to 89%, with a precision of 84%. For the dataset including human papillomavirus history, the support vector machine model achieved the highest balanced performance, yielding similar metrics. However, because these models were developed and evaluated strictly within a single institution, these balanced metrics represent optimistic internal validation performance.
Conclusions:
Machine learning models show promising performance for predicting cervical intra-epithelial neoplasia grade 2 or higher, with improved sensitivity after addressing class imbalance. These findings support further external, multi-center validation before clinical implementation.