Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Temporal evolution of large language models (LLMs) in oncology
Zilin Qiu1, Aimin Jiang2, Chang Qi3
1Department of Oncology, Zhujiang Hospital, Southern Medical University, Donghai County People's Hospital (Affiliated Kangda College of Nanjing Medical University), Lianyungang, 222000, China.
Background:
Large language models (LLMs) are increasingly being applied in healthcare; however, their performance in specialized fields, such as oncology, is subject to temporal factors, including knowledge decay and concept drift. The impact of these temporal dynamics on LLM question-answering accuracy in oncology remains inadequately evaluated. This study aims to systematically assess the temporal evolution of LLM accuracy in responding to oncology-related questions using real-world data.
Method:
We systematically collected relevant literature through 2025 by searching LLM-related keywords in PubMed, Google Scholar, and Web of Science databases. The inclusion criteria were as follows: (1) cancer-related research; (2) clear and complete question descriptions; and (3) complete answers. The final sample (n = 23) contained 614 research questions, comprising subjective questions (n = 223) and multiple-choice questions (n = 391). Following randomization of responses generated by three LLMs (ChatGPT-3.5, ChatGPT-4, and Gemini), we evaluated their accuracy across different cancer categories using both original scoring criteria and Likert scale scoring methods. Data analysis was performed using R statistical software, employing random or fixed effects models to calculate pooled mean differences (MD) and relative risks (RR) with their 95% confidence intervals (CI).
Results:
The findings demonstrated that in both subjective and objective oncology assessments, ChatGPT-3.5 (subjective questions MD = -3.30; objective questions RR = 0.92) and ChatGPT-4 (subjective questions MD = -7.17; objective questions RR = 0.93) showed declining performance trends over time, while Gemini exhibited significant improvements over time (subjective questions MD = 11.48; objective questions RR = 1.15). Notably, ChatGPT-3.5's performance on subjective questions revealed a significant turning point between March 14, 2023, and April 26, 2023, shifting from initially superior performance on newer questions to inferior performance compared with original questions, with the performance gap progressively widening.
Conclusions:
Our meta-analysis reveals temporal performance degradation in ChatGPT-3.5 and ChatGPT-4, which contrasts with the consistent improvement observed in Gemini. These findings provide essential guidance for the evidence-based deployment of LLMs in oncology.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Tumor Progression
Colon cancer is one of the best-documented examples of tumor progression. Early mutation in the APC gene in colon cells causes a small growth on the colon wall called a polyp. With time, this polyp grows into a benign, pre-cancerous tumor. Further...
Cancer Survival Analysis
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...

