Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Temporal evolution of large language models (LLMs) in oncology
Zilin Qiu1, Aimin Jiang2, Chang Qi3
1Department of Oncology, Zhujiang Hospital, Southern Medical University, Donghai County People's Hospital (Affiliated Kangda College of Nanjing Medical University), Lianyungang, 222000, China.
Large language models (LLMs) in oncology show performance decline over time, with Gemini improving while ChatGPT models degrade. This impacts evidence-based LLM deployment in cancer care.
Area of Science:
- Artificial Intelligence in Oncology
- Natural Language Processing in Medicine
- Machine Learning for Healthcare
Background:
- Large language models (LLMs) are increasingly used in healthcare, but their accuracy in specialized fields like oncology can be affected by knowledge decay and concept drift.
- The temporal dynamics influencing LLM performance in oncology question-answering require systematic evaluation.
- This study assesses the evolution of LLM accuracy for oncology-related queries using real-world data.
Purpose of the Study:
- To systematically evaluate the temporal performance trends of different LLMs in answering oncology-related questions.
- To compare the accuracy evolution of ChatGPT-3.5, ChatGPT-4, and Gemini in oncology.
- To provide insights for the evidence-based deployment of LLMs in clinical oncology settings.
Main Methods:
- A systematic literature search was conducted through 2025 using keywords related to LLMs and cancer research.
- A dataset of 614 oncology research questions (subjective and multiple-choice) was curated.
- Accuracy of responses from ChatGPT-3.5, ChatGPT-4, and Gemini was evaluated over time using standardized scoring methods and statistical analysis (random/fixed effects models).
Main Results:
- ChatGPT-3.5 and ChatGPT-4 demonstrated declining performance in both subjective and objective oncology question assessments over time.
- Gemini showed significant performance improvements over the study period for both question types.
- ChatGPT-3.5 experienced a notable performance shift, degrading over time, particularly for subjective oncology questions.
Conclusions:
- Meta-analysis indicates temporal performance degradation in ChatGPT-3.5 and ChatGPT-4 within the oncology domain.
- Gemini exhibited consistent performance improvement over time, contrasting with the degradation observed in other models.
- Findings offer crucial guidance for the judicious and evidence-based integration of LLMs into oncology practice.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Tumor Progression
Colon cancer is one of the best-documented examples of tumor progression. Early mutation in the APC gene in colon cells causes a small growth on the colon wall called a polyp. With time, this polyp grows into a benign, pre-cancerous tumor. Further...
Cancer Survival Analysis
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...

