Related Experiment Video
Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Evaluation of a Medical Large Language Model in Answering Real-World Radiation Oncology Questions:
Fabio Dennstädt1, Max Schmerder1, Elena Riggenbach1
1Inselspital, Department of Radiation Oncology, Bern University Hospital, University of Bern, Bern, Switzerland.
A locally deployed large language model (LLM) demonstrated comparable performance to clinical experts in answering radiation oncology questions. This medical AI shows potential for clinical support, though further evaluation is needed for widespread adoption.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Radiation Oncology Applications
Background:
- Large language models (LLMs) show promise for clinical tasks in specialized fields like radiation oncology.
- Previous LLM evaluations focused on exam settings, leaving real-world clinical performance uncertain.
- The potential of locally deployed LLMs as AI assistants for clinical questions remains to be determined.
Purpose of the Study:
- To evaluate a state-of-the-art medical LLM's performance against clinical experts in answering real-world radiation oncology questions.
- To assess the quality and potential harmfulness of LLM-generated answers for clinical decision-making.
Main Methods:
- Physicians collected clinical radiation oncology questions from 10 European hospitals.
- Fifty questions were answered by 3 senior radiation oncology experts and the LLM Llama3-OpenBioLLM-70B.
- Physicians conducted a blinded review, rating answer quality, potential harmfulness, and recognizability of the source (expert vs. LLM).
Main Results:
- No significant difference in answer quality was found between the LLM and clinical experts (mean scores 3.38 vs 3.63).
- Potentially harmful answers occurred in 16% of LLM responses versus 13% for clinical experts (not statistically significant).
- Physicians correctly identified the source of answers 78% of the time for experts and 72% for the LLM.
Conclusions:
- A locally deployed medical LLM performs comparably to clinical experts in answering radiation oncology questions regarding quality and safety.
- LLMs are becoming increasingly capable and affordable for hospital deployment, though not yet ready for full clinical implementation as general assistants.
- Real-world evaluation studies are crucial for understanding LLM limitations and guiding responsible clinical integration, necessitating healthcare professional education on generative AI.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Cancer Survival Analysis
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Comparing the Survival Analysis of Two or More Groups