Related Experiment Video
Updated: Jun 6, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
504
Assessment of Large Language Models (LLMs) in decision-making support for gynecologic oncology
Khanisyah Erza Gumilar1,2, Birama R Indraprasta3, Ach Salman Faridzi3
1Graduate Institute of Biomedical Science, China Medical University, Taichung, Taiwan.
Computational and Structural Biotechnology Journal
|November 29, 2024
Summary
Gemini Advanced (GemAdv) shows strong potential in gynecologic oncology, offering accurate and consistent answers for complex cancer cases. Further development is needed for optimal clinical integration.
Area of Science:
- Artificial Intelligence in Medicine
- Oncology Decision Support
- Natural Language Processing in Healthcare
Background:
- Large Language Models (LLMs) are rapidly advancing, necessitating rigorous evaluation for safe clinical integration.
- Assessing LLM reliability and accuracy is crucial for supporting medical professionals in complex casework.
- Ensuring LLM performance in specialized fields like gynecologic oncology is vital for their adoption.
Purpose of the Study:
- To investigate the accuracy and consistency of prominent Large Language Models (LLMs) in complex gynecologic cancer cases.
- To evaluate the performance of ChatGPT-4, Gemini Advanced, and Copilot in a clinical decision-making context.
- To determine the suitability of LLMs for supporting gynecologic oncologists.
Main Methods:
- Three LLMs (ChatGPT-4, Gemini Advanced, Copilot) were assessed using 15 clinical vignettes and 5 open-ended questions.
- Responses were evaluated blindly by six expert gynecologic oncologists on a 5-point Likert scale.
- Metrics included accuracy, consistency, relevance, clarity, depth, focus, and coherence.
Main Results:
- Gemini Advanced demonstrated superior accuracy (81.87%) compared to ChatGPT-4 (61.60%) and Copilot (70.67%).
- Gemini Advanced consistently provided correct answers more frequently (>60% daily).
- While ChatGPT-4 showed a slight advantage in NCCN guideline adherence, Gemini Advanced excelled in answer depth and focus.
Conclusions:
- LLMs, particularly Gemini Advanced, show promise for supporting clinical practice in gynecologic oncology with accurate and consistent information.
- Further refinement of LLMs is necessary for handling highly complex gynecologic cancer scenarios.
- Ongoing development and rigorous evaluation are essential to maximize the clinical utility and reliability of LLMs in oncology.

