Performance of clinical breast cancer risk prediction models vs a mammography-based artificial intelligence risk
Medha Kaul1, Christopher G Scott1, Alena Wadzinske1,2
1Department of Quantitative Health Sciences, Mayo Clinic, Rochester, MN, United States.
Background:
Artificial intelligence (AI)-based, mammography breast cancer risk prediction models show improved discriminatory accuracy relative to clinical risk models. However, data on their calibration are limited. This study compared model performance of 3 clinical breast cancer risk models: Gail, Tyrer-Cuzick v8, and Breast Cancer Surveillance Consortium v3 with the MIRAI AI-risk model.
Methods:
Digital mammograms were ascertained from a screening mammography cohort of 12 308 women within the Mayo Clinic Biobank with 250 incident breast cancers (176 invasive) within 5 years. We predicted 5-year breast cancer risk, estimated discriminatory accuracy (C-index), and calibration (observed-to-expected ratio) of both overall and invasive breast cancer and compared estimates using bootstrapping approaches.
Results:
MIRAI demonstrated similar or improved discriminatory accuracy of overall breast cancer (C-index = 0.71, 95% confidence interval [CI] = 0.68 to 0.74) and invasive breast cancer (C-index = 0.71, 95% CI = 0.67 to 0.75) compared with clinical models (overall breast cancer: C-index = 0.59-0.68; invasive breast cancer: C-index = 0.60-0.68). MIRAI's calibration for risk of overall breast cancer (observed-to-expected ratio = 0.96, 95% CI = 0.85 to 1.08) was improved compared with Gail (observed-to-expected ratio = 1.22, 95% CI = 1.07 to 1.38) and Breast Cancer Surveillance Consortium (observed-to-expected ratio = 1.38, 95% CI = 1.22 to 1.56) but similar to TC with volumetric percent density and polygenic risk score (observed-to-expected ratio = 0.99, 95% CI = 0.87 to 1.13). However, for low-risk women (approximately 50%), MIRAI overestimated risk of overall breast cancer. MIRAI also overestimated risk of invasive breast cancer across the risk spectrum (observed-to-expected ratio = 0.68, 95% CI = 0.58 to 0.78), whereas clinical models had good calibration (observed-to-expected ratio = 0.86-0.99).
Conclusion:
MIRAI demonstrated stronger discriminatory accuracy than clinical models for 5-year overall and invasive breast cancer risk prediction but overestimated risk for both breast cancer endpoints. AI-based risk models should consider discriminatory accuracy and calibration for invasive cancer before implementation.


