Related Experiment Video
Updated: Jun 19, 2026

Universal Screening for Prevention of Reading, Writing, and Math Disabilities in Spanish
Published on: July 18, 2020
Collaborative large language models (LLMs) are all you need for screening in systematic reviews
Mihir Parmar1,2, Syed Arsalan Ahmed Naqvi1, Kainat Warraich1
1Division of Hematology and Oncology, Department of Medicine, Mayo Clinic, Phoenix, AZ.
Background:
The ability of large language models (LLMs) to work collaboratively and screen studies in a systematic review (SR) is under-explored. Hence, we aimed to evaluate the effectiveness of LLMs in automating the process of screening in systematic reviews.
Methods:
This is an observational study which included labeled data (title and abstracts) for five SRs. Originally, two reviewers screened the citations independently for eligibility. A third reviewer cross-checked each citation for quality assurance. GPT-4, Claude-3-Sonnet, and Gemini-Pro-1.0 were used using zero-shot chain-of-thought prompting. Collaborative approaches included (i): conflict resolution using benefit of the doubt, (ii) majority voting using an independent third LLM and (iii) conflict resolution using an informed third LLM. Performance was assessed using accuracy, precision for exclusion, and recall for inclusion. Work saved over samples (WSS) was computed to estimate the reduction in manual human effort.
Results:
A total of 11300 articles were included in this study. The individual models, GPT-4, Claude-3-Sonnet, and Gemini-Pro-1.0 exhibited a high precision for exclusion, achieving 99.7%, 99.7%, and 99.2% and high recall for inclusion achieving 95.5%, 96.6% and 85.7%, respectively. However, the collaborative approach utilizing the two best-performing models (GPT-4 and Claude-3S) achieved an average precision of 99.9% and a recall of 98.5% (across all collaborative approaches). Furthermore, the proposed collaborative approach resulted in an average WSS of 63.5%, compared to the average WSS of 45.2% for individual models. Conversational LLM interactions showed a consistent pattern of results.
Limitations:
This study was limited due to reliance on proprietary models, and evaluation on oncology datasets.
Conclusion:
Evidence shows that collaborative LLMs enable efficient, high-performing screening in systematic reviews, supporting continuous evidence updates.
Primary Funding Source:
NIH (U24CA265879-01-1) and Carolyn-Ann-Kennedy-Bacon Fund.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Stereotype Content Model
Models, Theories, and Laws
Mechanistic Models: Compartment Models in Individual and Population Analysis
Typical Model Studies
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language