Related Experiment Video
Updated: Sep 20, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
LLM-Based Classifiers Can Reduce the Proportion of Records to Screen with Minimal Loss of Relevant Studies: A Case
Barbara Nussbaumer-Streit1, Andreea Dobrescu1, Amin Sharifan1
1Department for Evidence-based Medicine and Evaluation, University for Continuing Education Krems, Austria.
Background:
The Epistemonikos Sustainable Knowledge Platform (SKP) integrates a suite of machine learning tools designed to make evidence synthesis more efficient, including access to the Epistemonikos Database of Trials (ED-Trials). One feature is the use of classifiers combined with large language models (LLMs) to support abstract screening during study selection. LLM-based classifiers can be applied to populations, interventions, and study design features. Their primary aim is to reduce the proportion of records needed to be screenedby automatically excluding records with a high probability of being ineligible. We aimed to evaluate SKP's LLM-based classifiers on abstract screenings of records retrieved via: (scenario 1) a traditional systematic literature search (original ), and (scenario 2) a search conducted within ED-Trials and subsequently forwarded to SKP.
Methods:
We used data from a completed but at the time unpublished systematic review and network meta-analysis (NMA)as reference standards. We evaluated 29 LLM-based classifier combinations in both scenarios and assessed the proportion of records excluded by automated screening as well as whether any relevant study was wrongly excluded and therefore missed. To determine the impact of missed studies, we re-ran the network meta-analyses of the original review, excluding missed studies from the model and we evaluated how missed studies impact the certainty-of-evidence ratings.
Results:
LLM-based classifiers excluded by automated screening 5.8-46.6% of 1,129 records identified by the original search and 7.7-34.3% of 1,003 records based on the SKP search. More specific classifiers excluded more records than broader ones. Of the 29 combinations applied to the original search results, 10 incorrectly excluded the same two relevant studies. Based on the SKP review, four combinations missed one study, and 10 combinations missed two studies. However, exclusion of these studies resulted in minimal changes to NMA-effect estimates (risk ratio changes of 0.0-0.04 in both scenarios) and did not alter certainty-of-evidence ratings.
Discussion:
LLM-based classifiers are a promising strategy for reducing the amount of records to screen while keeping the risk of missing relevant studies low. However, further evaluations using multiple-use cases across different medical topics are necessary to learn more about the generalizability of our results.
More Related Videos
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025