Related Experiment Videos
Machine learning for population-level risk prediction of future cholangiocarcinoma
Felix van Haag1, Jan Clusmann2, Paul-Henry Koop3
1Department of Internal Medicine III, Gastroenterology, Metabolic Diseases and Intensive Care, University Hospital RWTH Aachen, Aachen, Germany.
Ebiomedicine
|August 12, 2026
Summary
Machine learning models using routinely available clinical data can effectively stratify cholangiocarcinoma (CCA) risk. This approach offers a valuable tool for early detection and personalized screening of this aggressive cancer.
Area of Science:
- Oncology
- Medical Informatics
- Machine Learning
Background:
- Cholangiocarcinoma (CCA) has a poor prognosis due to late diagnosis, as current screening strategies are lacking.
- An early, cost-effective, and widely applicable risk assessment for CCA is needed.
Purpose of the Study:
- To develop and validate machine learning (ML) models for early risk stratification of cholangiocarcinoma (CCA).
- To identify key clinical predictors for CCA risk assessment.
Main Methods:
- Developed ML models using multimodal data from 487,495 UK Biobank participants.
- Reduced model inputs to five and ten routinely available clinical parameters.
- Externally validated models in four independent international cohorts (PMBB, AOU, JMDC, TriNetX).
Main Results:
- ML models integrating health records and Gamma glutamyltransferase demonstrated robust CCA risk stratification.
- Achieved good performance across diverse cohorts with AUROCs ranging from 0.71 to 0.8.
- Identified a number needed to screen of 79 in the All of Us Research Program cohort.
Conclusions:
- A comprehensive framework for early CCA risk stratification in the general population was established.
- Demonstrated the potential of data-driven models for personalized screening of hepatobiliary cancers.
- Key predictors for CCA risk were identified, facilitating targeted screening efforts.