Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models in Automated Medical Literature Screening: A Systematic Review and Meta-Analysis
Chenggong Xie1, Weichang Kong2, Liuting Pi3
1Institute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, Beijing, China.
Journal of Evidence-Based Medicine
|July 25, 2026
Summary
Large language models (LLMs) demonstrate high accuracy in automated medical literature screening, especially for full-text assessment. These AI tools significantly reduce manual screening workload, aiding evidence synthesis.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Systematic Review Methodology
Background:
- Automated screening of medical literature is crucial for efficient evidence synthesis.
- Large Language Models (LLMs) offer potential for enhancing literature screening processes.
Purpose of the Study:
- To systematically evaluate the diagnostic performance of LLMs in automated medical literature screening.
- To assess the potential role of LLMs in supporting evidence synthesis workflows.
Main Methods:
- A systematic review and meta-analysis of studies evaluating LLMs for medical literature screening (title/abstract and full-text).
- Searched multiple databases (PubMed, Web of Science, Embase, etc.) from January 2022 to June 2026.
- Pooled sensitivity, specificity, and AUC were calculated using bivariate random-effects models.
Main Results:
- Pooled sensitivity for title/abstract screening was 0.92 and specificity was 0.94 (AUC 0.98).
- Pooled sensitivity and specificity for full-text screening reached 0.99 (AUC 0.99).
- LLMs demonstrated significant efficiency gains, reducing workload and screening time substantially.
Conclusions:
- LLMs show promising diagnostic performance for automated medical literature screening, particularly for full-text assessment.
- LLMs can serve as effective assistive tools, significantly reducing the burden of manual screening.
- Further real-world validation is recommended to integrate LLMs into evidence-based medicine practices.
