Related Experiment Video
Updated: May 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models in neuro-oncology: A comparative study of accuracy, completeness, and clinical
Shefqet Hajdari1, Minaam Farooq2, Aleeza Habib2
1University Hospital Bonn, Department of Neurosurgery, Bonn, Germany.
Background:
Large language models (LLMs), with their remarkable ability to retrieve and analyse the information within seconds, are generating significant interest in the domain of healthcare. This study aims to assess and compare the accuracy, completeness, and usefulness of the responses of Gemini Advanced, ChatGPT-3.5, and ChatGPT-4, in neuro-oncology cases.
Methods:
For 20 common neuro-oncology cases, four questions regarding differential diagnosis, diagnostic workup, provisional diagnosis and management plan were asked from 3 LLMs. All the responses after replication and blinding were evaluated by senior neuro-oncologists on 3 scales: Accuracy (1-6), completeness (1-3), and usefulness (1-3). To compare the performance of all three LLMs, ANOVA and Dunn's post hoc test were employed. A p-value of less than 0.05 was considered statistically significant.
Results:
Combining all domains, ChatGPT-4 remained the most accurate (x̅ = 5.12), the most complete (x̅ = 2.05) and the most useful (x̅ = 2.06) followed by Gemini Advanced (x̅ = 4.97, 2.04, 2.05) and ChatGPT-3.5 (x̅ = 4.9, 1.9, 1.99). Using Dunn's post hoc test, completeness of differential diagnosis has two statistically significant pairs, ChatGPT-3.5 and ChatGPT-4 (p = 0.001), and ChatGPT-3.5 and Gemini Advanced (p = 0.003). For overall completeness, statistically significant difference was found between three LLMs according to ANOVA (p < 0.001).
Conclusion:
For further studies, more extensive tests should be conducted with a larger number of neurosurgeons assessing the responses to a more diverse range of clinical cases to better understand their strengths and limitations. The performance of Retrieval Augmented Generation (RAG), after specific training of LLMs for treatment guidelines of neuro-oncology cases can also be assessed in future studies.
More Related Videos
Related Concept Videos
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Aneurysm II: Clinical Manifestations and Diagnostic Studies

