Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A multidimensional benchmarking framework for large language models in oncologic decision making
Mehmet Halici1, Serkan Salturk2, Irem Sayin3
1Department of Radiation Oncology, Basaksehir Cam and Sakura City Hospital, 34480, Istanbul, Turkey. mehmethalici95@gmail.com.
Scientific Reports
|July 21, 2026
Summary
Evaluating large language models (LLMs) for oncology clinical decisions requires a multi-dimensional approach. GPT-5 excelled in accuracy and efficiency, demonstrating the value of comprehensive performance scoring for AI tools in cancer care.
Area of Science:
- Artificial Intelligence in Oncology
- Clinical Decision Support Systems
- Natural Language Processing
Background:
- Large language models (LLMs) show promise for clinical decision support in oncology.
- Current evaluations often rely on single metrics, hindering comprehensive assessment.
- Developing multi-dimensional frameworks is crucial for effective LLM implementation.
Purpose of the Study:
- To compare the performance of leading LLMs (Gemini 2.5 Pro, GPT-5, Claude Opus 4.1) in realistic oncology scenarios.
- To establish a multi-dimensional evaluation framework integrating clinical accuracy, explainability, and operational efficiency.
- To determine the optimal LLM for clinical decision support in non-small cell lung cancer (NSCLC).
Main Methods:
- A comparative observational study using five stepwise NSCLC clinical scenarios.
- LLMs answered open-ended clinical questions via official APIs, compared against evidence-based answers.
- Model outputs were assessed for clinical accuracy, explainability, cost, response time, and generative efficiency.
- A Composite Performance Score (CPS) integrated all evaluated dimensions.
Main Results:
- Significant inter-model differences were observed across all evaluated metrics (p < 0.001).
- GPT-5 led in accuracy, explainability, and generative efficiency.
- Gemini 2.5 Pro offered the lowest cost; Claude Opus 4.1 had the fastest response times.
- GPT-5 achieved the highest Composite Performance Score, followed by Gemini 2.5 Pro and Claude Opus 4.1.
Conclusions:
- A multi-dimensional evaluation framework provides more actionable insights than single-metric assessments for LLM selection in oncology.
- GPT-5 demonstrated superior overall performance in the evaluated NSCLC clinical scenarios.
- Clinician supervision remains essential when utilizing LLMs for oncology decision support.