Predicting Early-Onset Colorectal Cancer with Large Language Models
Wilson Lau1, Youngwon Kim1, Sravanthi Parasa2
1Truveta, Bellevue, WA.
AMIA ... Annual Symposium Proceedings. AMIA Symposium
|February 23, 2026
Summary
Early-onset colorectal cancer (EoCRC) is rising in younger adults. A fine-tuned large language model (LLM) shows promise in predicting EoCRC using patient data, achieving 73% sensitivity and 91% specificity.
Area of Science:
- Oncology
- Artificial Intelligence
- Medical Informatics
Background:
- Early-onset colorectal cancer (EoCRC) incidence is increasing in individuals under 45.
- This demographic is below current national cancer screening age guidelines.
- Predictive models are needed to identify at-risk individuals for EoCRC.
Purpose of the Study:
- To evaluate machine learning (ML) and large language models (LLMs) for predicting early-onset colorectal cancer (EoCRC).
- To compare the predictive performance of various ML models against advanced LLMs.
- To utilize patient journey data within six months preceding diagnosis for EoCRC prediction.
Main Methods:
- Retrospective analysis of 1,953 colorectal cancer (CRC) patients from US health systems.
- Application and comparison of 10 distinct machine learning models.
- Utilized a fine-tuned large language model (LLM) incorporating patient conditions, lab results, and observations.
Main Results:
- The fine-tuned LLM demonstrated superior performance in predicting EoCRC.
- Achieved an average sensitivity of 73% for EoCRC prediction.
- Achieved an average specificity of 91% for EoCRC prediction.
Conclusions:
- Fine-tuned LLMs show significant potential for early detection of colorectal cancer in younger populations.
- LLMs can effectively analyze diverse patient data for cancer risk prediction.
- This approach may aid in developing targeted screening strategies for early-onset colorectal cancer.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Improving Translational Accuracy
3.7K
3.7K


