Related Experiment Video
Updated: Jun 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Applications of large language models in tumor boards: a systematic review
Benjamin Konzmann1, Maurice Henkel2, Christian Breit3
1Institute of Urology, University Hospital Basel, Basel, Switzerland.
Background:
Multidisciplinary tumor boards (MDTs) are the gold standard for cancer care but currently face significant pressure from rising case volumes and the increasing complexity of precision medicine. Large Language Models (LLMs) offer potential as Clinical Decision Support Systems (CDSS) to augment these workflows. This systematic review evaluates the current applications, accuracy, and safety of LLMs in MDT decision-making.
Methods:
A systematic review was conducted following PRISMA 2020 guidelines. We searched PubMed/MEDLINE, Embase, IEEE Xplore, and arXiv for primary research published between January 2023 and November 2025. Inclusion criteria required studies to evaluate LLM performance in tumor board settings against human consensus or established clinical guidelines.
Results:
Thirty-one studies encompassing 3,845 unique patient cases were included. The analysis revealed a distinct "complexity gap" in model performance. In standardized, high-incidence domains such as breast and prostate cancer, advanced models (particularly the GPT-4 family) demonstrated high concordance (up to 94%) with human experts. However, performance degraded significantly in complex, rare, or multimodal scenarios-such as sarcoma or neuro-oncology-where models struggled with "gray zone" decision-making and lacked the ability to interpret non-textual data. While LLMs showed utility in administrative tasks and guideline retrieval, they remained prone to hallucinations and lacked the nuance required for holistic patient assessment.
Conclusion:
Current LLMs exhibit sufficient maturity to function as assistive tools for documentation and decision support in routine oncological cases but are not yet reliable enough for autonomous decision-making. Successful clinical implementation will require "human-in-the-loop" safeguards, the development of multimodal architectures, and rigorous prospective validation to ensure patient safety.
