Related Experiment Video
Updated: Jul 15, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models to Extract Cancer Staging Data From Clinical Documentation at Scale
Swapna Abhyankar1, Rajesh M Rao1, Mehraveh Salehi1
1Truveta, Inc, Bellevue, WA.
JCO Clinical Cancer Informatics
|July 13, 2026
Summary
A new large language model (LLM), TLM-Oncology, precisely extracts oncology staging data from clinical notes for multiple cancer types. This advances the use of previously inaccessible real-world data for cancer research.
Area of Science:
- Oncology
- Medical Informatics
- Natural Language Processing
Background:
- Extracting accurate cancer staging data from clinical documentation is crucial for patient care and research.
- Manual extraction is time-consuming and prone to errors, limiting the use of real-world data.
Purpose of the Study:
- To develop and evaluate the Truveta Language Model Oncology (TLM-Oncology), a large language model (LLM), for precise extraction of real-world oncology staging data.
- To assess the model's performance across multiple cancer types using clinical documentation.
Main Methods:
- A pretrained LLM was fine-tuned using supervised learning on annotated clinical notes for bladder, cervical, colorectal, breast, and prostate cancers.
- Performance was evaluated using precision, recall, and F1 scores at both relation and attribute levels.
Main Results:
- TLM-Oncology extracted over 2.5 million staging records for 217,768 patients from over two million notes.
- High relation-level precision (0.77-1.0) was achieved across multiple cancer types, demonstrating the model's effectiveness.
Conclusions:
- TLM-Oncology successfully extracts detailed cancer staging information from diverse clinical documentation with high precision.
- The model transforms previously inaccessible data into a valuable resource for downstream applications in oncology research and care.