Related Experiment Video
Updated: Aug 22, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Augmenting Head and Neck Multidisciplinary Tumor Board Recommendations With Locally Run Large Language Models:
Christoph Raphael Buhr1,2, Lukas Müller3, Daniel Pinto Dos Santos3
1Department of Otorhinolaryngology, University Medical Center of the Johannes Gutenberg-University Mainz, Langenbeckstraße 1, Mainz, Germany, 49 6131 17 7362.
Background:
Multidisciplinary tumor boards (MDTs) constitute the foundation of modern tumor therapy. Large language models (LLMs) are widely discussed for optimizing their recommendations.
Objective:
This is the first prospective feasibility study evaluating the implementation of locally run LLMs on real-world cases within a regular head and neck MDT.
Methods:
Seventeen patients participated in the study. The MDT cases were processed by 2 different local LLMs (gemma-3-12b and gpt-oss-20b) to obtain treatment recommendations. The MDT conferred as usual. After the decision was made, the MDT was presented with the LLMs' recommendations. If deemed to be beneficial, the MDT's recommendation was adjusted. The MDT members rated the LLMs' responses inter alia, for medical adequacy on a 6-point Likert scale. In addition, a tabular comparison of the MDT's and LLMs' recommendations was carried out.
Results:
In one case, 6% (1/17, 95% CI 0%-29%), the LLM was able to substantially improve the MDT recommendation by underscoring a follow-up examination that had not yet been performed. Concordance regarding the curative or palliative therapy regimen reached 94% (16/17, 95% CI 71%-100%); for gemma-3-12b and 59% (10/17, 95% CI 33%-82%) for gpt-oss-20b. Gemma-3-12b stated the same first-line therapy regimen as the MDT as first-line in 35% (6/17, 95% CI 14%-62%) of cases, and gpt-oss-20b in 41% (7/17, 95% CI 18%-67%) of cases. In 59% (10/17, 95% CI 33%-82%) of patients, gemma-3-12b stated the MDT's first-line therapy regimen, albeit with a different priority, while for gpt-oss-20b, it was 41% (7/17, 95% CI 18%-67%) of patients. Medical adequacy, as rated by the MDT members, revealed a median of 5 (IQR 2-5) for gemma-3-12b and 4 (IQR 3-5) for gpt-oss-20b. MDT members stated potentially hazardous information in 27% (25/93, 95% CI 18%-37%) of ratings for gemma-3-12b and 17% (14/83, 95% CI 9%-26%) of ratings for gpt-oss-20b.
Conclusions:
Locally run LLMs improved the MDT recommendation in 1 case and were primarily useful for identifying potentially relevant missing information in other cases, underscoring that they cannot replace MDTs. However, their observed benefit suggests that more advanced local models may offer safe, rapid, and cost-effective support for MDT decision-making. The study should be seen as an exploratory setting focusing on practical insights rather than benchmarking its clinical impact. Accordingly, the data demonstrate that the integration of LLMs in today's MDT workflow is feasible and may benefit the quality of decision-making in specific cases.