Related Experiment Video
Updated: Jun 19, 2026

A Mouse Model of Incompletely Resected Soft Tissue Sarcoma for Testing Neoadjuvant Therapies
Published on: July 28, 2020
When AI joins the table: evaluating large language model performance in soft tissue sarcoma tumor board decisions.
Reza Dehdab1, Saif Afat1, Fiona Mankertz1
1Department of Radiology, Tübingen University Hospital, University of Tübingen, Tübingen, Germany.
Large language models like ChatGPT-4o show promise in assisting multidisciplinary tumor boards for soft tissue sarcoma (STS) treatment recommendations, particularly in clinical contextualization. However, expert oversight remains crucial for treatment sequencing and chemotherapy selection.
Area of Science:
- Oncology
- Artificial Intelligence in Medicine
- Medical Informatics
Background:
- Multidisciplinary tumor boards (MDTs) are essential for personalized soft tissue sarcoma (STS) management.
- Current MDTs face limitations including time, cost, and resource constraints.
- Large language models (LLMs) present a potential avenue for augmenting MDT workflows.
Purpose of the Study:
- To evaluate the clinical performance of ChatGPT-4o in generating treatment recommendations for real-world STS cases.
- To compare ChatGPT-4o's suggestions against expert MDT decisions using predefined criteria.
- To assess LLM performance across different domains of cancer care and sarcoma subtypes.
Main Methods:
- Retrospective analysis of 152 STS patients' anonymized registration letters.
- ChatGPT-4o generated guideline-based treatment recommendations.
- Expert reviewers blinded to the source scored outputs across five domains: diagnostics, therapeutics, sequencing/timing, chemotherapy, and clinical contextualization.
Main Results:
- ChatGPT-4o performance scores were significantly below the maximum across all evaluated criteria (p < 0.0001).
- Clinical contextualization domain received significantly higher scores compared to other domains (p < 0.05).
- No significant performance differences were noted across various sarcoma subtypes (p = 0.138).
Conclusions:
- ChatGPT-4o demonstrated considerable expert-rated performance in generating STS tumor board recommendations, especially in clinical contextualization.
- Areas requiring improvement include treatment sequencing and chemotherapy selection, underscoring the need for expert human oversight.
- Findings support the integration of LLMs into oncology workflows, with further development needed for safe clinical application.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Related Concept Videos
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Tumor Immunotherapy