Related Experiment Video
Updated: Jun 1, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Automated evidence surveillance with AI-enabled preranking with cutoff in living guideline maintenance: a simulation
Darren Rajit1, Steve McDonald2, Lan Du3
1Monash Centre for Health Research and Implementation, Faculty of Medicine, Nursing, and Health Sciences, Monash University, Clayton, Victoria, Australia.
Objectives:
Living guidelines are an emerging approach to ensure timely synthesis of the research evidence. However, pragmatic methods for maintenance are needed to ensure sustainability. Our study aimed to simulate and evaluate the performance and efficiency of various single database evidence retrieval workflows augmented by artificial intelligence (AI)-enabled preranking with cutoff for living guideline development and maintenance.
Methods:
A retrospective simulation study was conducted using data from the 2023 International Polycystic Ovary Syndrome Guidelines. Simulations were run across four databases (Medline, Embase, PubMed, and OpenAlex) to identify the peer-reviewed articles included in the guidelines. Workflows were evaluated at the guideline (all articles) and topic level. Single database topic-specific searches were compared against single database overarching searches. The performance of overarching searches with AI-enabled preranking with cutoff at guideline and topic level was also evaluated. Metrics included recall, precision, F score, number of articles needed to screen per relevant study (NNR), and overall screening workload.
Results:
Across 38 eligible topics (854 articles), overarching searches outperformed topic-specific searches at guideline level for both recall (92%-96% vs 76%-89%) and efficiency, reducing overall screening workload by 63%-70% and requiring teams to screen 28-48 articles per relevant study vs 76-160 between comparable databases (Embase, Medline). At individual topic level, topic-specific searches were more efficient than overarching searches integrated with topic-specific rankings. However, topic-specific searches had significantly lower recall (P < .01) in comparison. AI-enabled ranking provided only marginal efficiency gains at guideline level (3%-21% NNR reduction) compared to topic level (85%-95% NNR reduction). Lastly, performance of automated article retrieval via PubMed application programming interface was equivalent to manual retrieval via Ovid Medline.
Conclusion:
Single database overarching searches outperform single database topic-specific searches and should be considered during guideline maintenance when most of the guideline needs updating. While topic-specific searches may be more efficient in instances where only a few areas need to be updated, using a single database approach may result in lower recall. Single database overarching searches integrated with topic-specific rankings can be considered in such cases.