Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
From Information Extraction to Clinical Reasoning: A Systematic Scoping Review of Large Language Models in Cancer
Maryam Seifaddini1, Mohammad Beheshti1, Steven Richberg2
1Missouri Cancer Registry and Research Center, Department of Public Health, College of Health Sciences, University of Missouri, Columbia, Missouri; MU Institute for Data Science and Informatics, University of Missouri, Columbia, Missouri.
Abstract:
Pathology reports anchor cancer diagnosis and staging, yet their narrative structure limits the reliable translation into structured, machine-actionable knowledge, creating a bottleneck between expert interpretation and scalable clinical intelligence. Despite decades of clinical natural language processing research, pathology text remains among the most complex and consequential sources of medical data to operationalize at scale. Large language models (LLMs) offer new approaches for reading, extracting, and interpreting these reports. We synthesized the current LLM applications in cancer pathology using a 4-level capability framework across the pathology-report data lifecycle: level 1, text preparation and quality checks; level 2, information extraction; level 3, guideline-based clinical reasoning, such as TNM staging and registry coding; and level 4, interpretive synthesis, such as explanations, summarization, or decision support. Rather than grouping studies by natural language processing task labels, this framework tracks how LLM applications progress from preprocessing and extraction toward higher-level interpretation and synthesis. We followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews guidelines and searched 4 databases through September 2, 2025, identifying 41 eligible studies. Most studies focused on level 2 tasks, with fewer addressing level 3 and level 4 tasks. Encoder-based models, including domain-specific variants, such as BioBERT, were commonly used for structured extraction tasks, whereas generative models, including GPT, LLaMA, and Mistral-family models, were increasingly evaluated for prompt-based extraction, staging, and summarization. Reported performance was often high for well-defined extraction tasks, but external validation was uncommon, and metrics varied across studies, limiting direct comparison. Overall, the evidence suggests that success at lower capability levels does not consistently translate into higher-level reasoning, especially when reports are inconsistent, required staging inputs are missing, or clinical assumptions must be inferred, which helps explain gaps between benchmark results and practical adoption. Future work should prioritize robust multisite validation, clinically meaningful error analysis, transparent evaluation, and privacy-preserving implementation strategies to support the safe integration of LLMs in oncology.
Related Concept Videos
Cancer Survival Analysis
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
