Related Experiment Video
Updated: Aug 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Precision oncology meets Generative AI: assessing large language models in multidisciplinary GIST tumor boards
Reza Dehdab1, Judith Herrmann1, Fiona Mankertz1
1Department of Radiology, University Hospital Tübingen, University of Tübingen, Tübingen, Germany.
BMC Cancer
|August 4, 2026
Summary
Large language models (LLMs) like ChatGPT-5 and Qwen3 show high agreement with expert multidisciplinary tumor board (MTB) recommendations for gastrointestinal stromal tumors (GISTs). While effective, diagnostic reasoning remains a challenge, highlighting the need for expert oversight in GIST management.
Area of Science:
- Oncology
- Artificial Intelligence
- Medical Informatics
Background:
- Gastrointestinal stromal tumors (GISTs) require complex, individualized management decisions.
- Multidisciplinary tumor boards (MTBs) are the standard of care for GISTs, but access is often limited.
- Evaluating the utility of artificial intelligence, specifically large language models (LLMs), in supporting MTB decision-making is crucial.
Purpose of the Study:
- To assess the performance of two LLMs (ChatGPT-5 and Qwen3) in generating GIST MTB recommendations.
- To compare LLM-generated recommendations against expert MTB decisions using standardized criteria.
- To identify specific domains where LLMs excel or struggle in replicating MTB consensus.
Main Methods:
- A retrospective analysis of 99 GIST cases discussed in an institutional MTB.
- Development of a structured prompt to extract clinical data and generate LLM treatment recommendations.
- Independent, blinded evaluation of LLM outputs (ChatGPT-5, Qwen3) by expert reviewers across five domains.
Main Results:
- Both LLMs demonstrated high concordance with expert MTB recommendations (mean normalized scores: ChatGPT-5=0.901, Qwen3=0.875).
- No significant performance difference was observed between ChatGPT-5 and Qwen3.
- Diagnostic recommendations were a weaker domain for both models compared to other evaluated aspects.
Conclusions:
- LLMs show significant potential as assistive tools within GIST MTB workflows.
- The performance of ChatGPT-5 and Qwen3 suggests they can effectively support GIST management recommendations.
- Expert oversight remains essential, particularly for complex diagnostic reasoning in GIST cases.
