Related Experiment Video
Updated: Jun 15, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
525
Evaluating the effectiveness of large language models in abstract screening: a comparative analysis.
Michael Li1, Jianping Sun2, Xianming Tan3,4
1Department of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, NC, 27599, USA.
Systematic Reviews
|August 21, 2024
Summary
Large language models (LLMs) show high potential for abstract screening in systematic reviews, achieving over 90% accuracy. While not replacing human experts, LLMs offer efficient, cost-effective AI-assisted workflows for meta-analysis.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Evidence Synthesis
Background:
- Systematic reviews and meta-analyses are crucial for evidence-based medicine.
- Abstract screening is a time-consuming bottleneck in systematic reviews.
- Evaluating AI tools for this task is essential for improving efficiency.
Purpose of the Study:
- To assess the performance of various large language models (LLMs) in abstract screening for systematic reviews.
- To compare LLM effectiveness, efficiency, and integration potential against human expert workflows.
- To identify the most promising LLMs for automating parts of the systematic review process.
Main Methods:
- Developed Python scripts to interface with LLM APIs (e.g., ChatGPT, Gemini, Llama, Claude).
- Evaluated LLM performance on three abstract databases using sensitivity, specificity, and accuracy metrics.
- Compared LLM screening results against human-curated inclusion decisions as the gold standard.
Main Results:
- LLM performance varied, with ChatGPT v4.0 achieving over 90% accuracy.
- LLMs demonstrated high sensitivity and specificity, indicating reliable screening capabilities.
- LLMs offer a cost-effective and efficient alternative to traditional manual abstract screening.
Conclusions:
- LLMs show significant promise for revolutionizing abstract screening in systematic reviews.
- They can function as autonomous AI reviewers or augment human expert workflows.
- LLMs are poised to reshape systematic review and meta-analysis processes, enhancing efficiency and potentially accuracy.

