Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models for Oncology Nursing Decision Support: Cross-Sectional Study
Qiongyu Zhou1, Yan Jia2, Haiqin Hu2
1School of Nursing, Zhejiang Chinese Medical University, Hangzhou, Zhejiang, China.
Journal of Medical Internet Research
|July 24, 2026
Summary
Large language models (LLMs) show promise in oncology nursing tasks, excelling in structured knowledge but needing improvement for complex clinical judgment. Their use is best suited for information retrieval and patient education, complementing professional expertise.
Area of Science:
- Artificial Intelligence in Healthcare
- Nursing Informatics
- Clinical Decision Support Systems
Background:
- Large language models (LLMs) are increasingly adopted in healthcare for clinical decision support and nursing education.
- Evidence regarding LLM performance in specialized nursing fields like oncology nursing is limited.
- Evaluating LLM performance is crucial due to the complexity and high-risk nature of oncology care.
Purpose of the Study:
- To compare the performance of various LLMs in oncology nursing decision support.
- To explore the applicability and limitations of LLMs in oncology nursing practice.
- To assess LLMs using standardized examination questions and case-based clinical scenarios.
Main Methods:
- Five LLMs (DeepSeek, Qwen, Spark-Desk, WiseDiag, ChatGPT) were evaluated.
- Tasks included case-based questions from clinical scenarios and standardized examination questions.
- Responses were rated by experienced oncology nurses on correctness, clarity, and conciseness; examination performance was assessed by accuracy and efficiency.
Main Results:
- LLMs demonstrated strong performance in structured knowledge and examination tasks, with accuracy rates from 77% to 93%.
- DeepSeek and ChatGPT completed tasks in a single interaction, while others required multiple interactions.
- Statistically significant differences were found among models, with DeepSeek outperforming ChatGPT in specific metrics.
Conclusions:
- LLMs are effective for information retrieval and knowledge organization in oncology nursing but limited in complex clinical judgment.
- LLM outputs should be used as supplementary information, interpreted with professional clinical judgment.
- Further research is needed to enhance LLM capabilities for dynamic, individualized oncology nursing care.