Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
From Information Extraction to Clinical Reasoning: A Systematic Scoping Review of Large Language Models in Cancer
Maryam Seifaddini1, Mohammad Beheshti1, Steven Richberg2
1Missouri Cancer Registry and Research Center, Department of Public Health, College of Health Sciences, University of Missouri, Missouri; MU Institute for Data Science and Informatics, University of Missouri, Missouri.
Large language models (LLMs) can process cancer pathology reports, but current applications primarily focus on basic extraction, with limited success in higher-level clinical reasoning and interpretation. Robust validation is needed for safe integration into oncology.
Area of Science:
- Oncology
- Medical Informatics
- Natural Language Processing
Background:
- Pathology reports are crucial for cancer diagnosis and staging but are difficult to translate into structured data.
- Existing clinical natural language processing (NLP) research has faced challenges in operationalizing pathology text at scale.
- Large language models (LLMs) present a promising avenue for improving the extraction and interpretation of pathology report data.
Purpose of the Study:
- To synthesize current research on LLM applications in cancer pathology.
- To evaluate LLM capabilities across a four-level framework of the pathology report data lifecycle.
- To identify challenges and future directions for LLM integration in oncology.
Main Methods:
- A systematic literature review following PRISMA-ScR guidelines was conducted.
- Searched four databases for studies on LLMs in cancer pathology up to September 2, 2025.
- Identified and analyzed 41 eligible studies based on a four-level capability framework (text preparation, information extraction, clinical reasoning, interpretive synthesis).
Main Results:
- Most studies focused on information extraction (level 2), with fewer studies addressing clinical reasoning (level 3) and interpretive synthesis (level 4).
- Encoder-based models (e.g., BioBERT) were common for extraction, while generative models (e.g., GPT, LLaMA) showed potential for staging and summarization.
- High performance on extraction tasks often did not translate to higher-level reasoning, particularly with inconsistent reports or missing clinical data.
Conclusions:
- LLM success in basic extraction does not guarantee proficiency in complex clinical reasoning or interpretation.
- Gaps between benchmark results and practical adoption stem from report inconsistencies and the need for inferred clinical assumptions.
- Future research must prioritize multi-site validation, clinically relevant error analysis, transparent evaluation, and privacy-preserving implementation for safe integration in oncology.
Related Concept Videos
Cancer Survival Analysis
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
