Related Experiment Video
Updated: Jun 17, 2026

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
Published on: May 22, 2018
Comparing the performance of ChatGPT, DeepSeek, and Gemini in systematic and umbrella review tasks over time
Background:
This study aimed to compare the performance of ChatGPT-4o (OpenAI), DeepSeek-V3 (High-Flyer), and Gemini 1.5 Pro (Google) during 3 consecutive weeks in performing full-text screening, data extraction, and risk of bias assessment tasks in systematic and umbrella reviews.
Methods:
This study evaluated the correctness of large language model (LLM) responses in performing review study tasks by prompting 3 independent accounts. This process was repeated during 3 consecutive weeks for 40 primary studies. The correctness of responses was scored, and data were analyzed by Kendall W, generalized estimating equations followed by pairwise comparisons with Bonferroni correction, and Mann-Whitney U tests (α = .05).
Results:
DeepSeek achieved the highest data extraction accuracy (> 90%), followed by ChatGPT (> 88%). Moreover, DeepSeek outperformed significantly in data extraction compared with Gemini in most pairwise comparisons (P < .0167). Gemini showed an improvement in data extraction performance over time, with significantly higher accuracy in the third week than in the first week (P < .0167). ChatGPT generally performed better in systematic reviews than in umbrella reviews (P < .05).
Conclusions:
The studied LLMs showed potential for accurate data extraction, particularly DeepSeek, but consistently had unreliable performance in critical tasks like full-text screening and risk of bias assessment. LLM applications in review studies require cautious expert supervision.
Practical Implications:
Researchers planning to use LLMs for review study tasks should be aware that LLM responses to full-text screening and risk of bias assessment are unreliable. DeepSeek is the preferred LLM for data extraction in both systematic and umbrella reviews, whereas ChatGPT is recommended for systematic reviews.
