Related Experiment Video
Updated: Jun 19, 2026

Universal Screening for Prevention of Reading, Writing, and Math Disabilities in Spanish
Published on: July 18, 2020
Collaborative large language models (LLMs) are all you need for screening in systematic reviews
Mihir Parmar1,2, Syed Arsalan Ahmed Naqvi1, Kainat Warraich1
1Division of Hematology and Oncology, Department of Medicine, Mayo Clinic, Phoenix, AZ.
Collaborative large language models (LLMs) significantly improve systematic review screening efficiency and performance. This approach saves substantial human effort, supporting continuous evidence updates.
Area of Science:
- Artificial Intelligence in Medical Research
- Natural Language Processing for Evidence Synthesis
Background:
- Systematic reviews (SRs) require rigorous study screening, a process often bottlenecked by manual effort.
- The potential of large language models (LLMs) to automate and enhance SR screening remains largely unexplored.
Purpose of the Study:
- To evaluate the effectiveness of LLMs in automating the screening process for systematic reviews.
- To compare the performance of individual LLMs versus collaborative LLM approaches.
Main Methods:
- An observational study using labeled data (titles and abstracts) from five SRs.
- Individual LLMs (GPT-4, Claude-3-Sonnet, Gemini-Pro-1.0) and collaborative LLM strategies were employed.
- Performance metrics included accuracy, precision for exclusion, recall for inclusion, and work saved over samples (WSS).
Main Results:
- Individual LLMs demonstrated high precision for exclusion (up to 99.7%) and recall for inclusion (up to 96.6%).
- Collaborative LLM approaches, particularly using GPT-4 and Claude-3S, achieved superior average precision (99.9%) and recall (98.5%).
- Collaborative LLMs resulted in an average WSS of 63.5%, significantly higher than individual models (45.2%).
Conclusions:
- Collaborative LLMs offer an efficient and high-performing solution for systematic review screening.
- This automation supports the timely and continuous updating of evidence synthesis.
- Future research should explore LLM capabilities across diverse datasets and proprietary models.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Stereotype Content Model
Models, Theories, and Laws
Mechanistic Models: Compartment Models in Individual and Population Analysis
Typical Model Studies
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language