Related Experiment Video
Updated: May 1, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Novel Abstract Screening Algorithm Using Delphi-Inspired Large Language Model Consensus for Systematic Reviews in
Mirkamal Tolend1, Ramzi Halabi2, Kousai Ghaouari1
1Campbell Family Research Institute, Centre for Addiction and Mental Health, Toronto, Ontario, Canada.
Abstract:
BackgroundLarge language models (LLMs) may reduce the burden associated with performing systematic reviews by prescreening abstracts from a literature search for eligibility for inclusion in full-text review.MethodsWe developed an iterative, LLM-based workflow for screening abstracts: after manual specification of eligibility criteria and seed examples, an ensemble of five LLMs deliberates through a Delphi process to classify a batch of abstracts; these labels are used to train a logistic regression model that ranks the remaining abstracts and identifies a new batch of abstracts for LLM escalation until all abstracts are labelled by the LLM or probability thresholds. We tested our workflow on abstracts screened in three published systematic reviews in psychiatry. Our primary endpoint was the recall metric, and secondary endpoint was the work saved over sampling at 95% recall metric (WSS@95%).ResultsIn a dataset on autism biomarkers, 1,655 (35%) of 4,745 retrieved abstracts were judged to be relevant by the original authors. The Delphi-LLM workflow correctly identified 1,605 (97.0%) of these 1,655 abstracts (precision = 54.2%, WSS@95% = 38.1%). The performance metrics were better than non-LLM approaches (recall ≤ 91%, WSS@95 ≤ 26%), and, overall, balanced these metrics optimally compared to single-LLM agents (recall = 84.9-99.9%, WSS@95% = 16.7-39.8%). The recall and work saved metrics were similarly reliable and among the top in two low-prevalence datasets on an attention-deficit hyperactivity disorder treatment review (10% of 2,891 relevant) and a posttraumatic stress disorder trajectory review (7% of 4,453 relevant). For these two datasets, recall was 100.0% and 96.4%, and the WSS@95% was 17.3% and 18.5%, respectively.ConclusionsWe presented the design and validation of a novel abstract screening workflow that centres around a Delphi-style aggregation process to harness the strengths of five open-source LLMs that can be run on consumer-level workstations. This multi-LLM workflow showed acceptable and reliable performance for use as an automated prescreening method to facilitate systematic reviews.
More Related Videos
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025