Related Experiment Video
Updated: Sep 30, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Human-in-the-loop machine learning with community engagement reduces title, abstract and full-text screening workload
Nancy B Tahmo1, Anthony Noah2, Byron Odhiambo3
1Division of Epidemiology, Dalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada; MAP Centre for Urban Health Solutions, St. Michael's Hospital, Unity Health Toronto, Toronto, Ontario, Canada.
Objective:
Manual screening of titles, abstracts, and full texts for large-volume knowledge synthesis requires significant time investment. Machine-assisted tools offer solutions for facilitating screening. Most applications to date have focused on actively reordering citations or suggesting screening decisions for reviewers to prioritize relevant studies, or on automating exclusions based on narrowly defined eligibility criteria. However, few tools extend to full-text screening, and among those that report, a majority have largely been developed with little engagement of knowledge users. Knowledge user input has the potential to foster the relevance of review outcomes and uncover nuances such as broader community participation in research, that is often variably defined and applied. We aimed to describe a pragmatic machine-assisted screening workflow that extends to full-text screening and engages knowledge users throughout.
Study Design And Setting:
We implemented a pragmatic machine-assisted workflow for screening in a large-volume scoping review of community participatory approaches in infectious disease mathematical modeling. The review considered studies on any specified population, modeling any infectious disease, and where community participation was reported in at least one modeling step. The workflow was co-developed with community researchers, representing community-based organizations, who contributed to defining eligibility criteria and screening. First, we screened titles and abstracts using a supervised random forest model. We then triaged full texts based on keywords and made final classifications (inclusion/exclusion) using ridge-penalized logistic regression. We evaluated the machine learning model performances using held-out test sets and bootstrapped 95% confidence intervals.
Results:
Of 16,629 studies identified in the search, 9,841 unique citations underwent title and abstract screening and 2786 citations underwent full-text screening. Model-based title and abstract screening achieved sensitivity of 86.7% (95% CI: 84.1-89.4) and area under the curve (AUC) of 86.4% (95% CI: 84.4-88.3), reducing screening workload by 48.3%. Model-based full-text screening achieved sensitivity of 82.6% (95% CI: 66.7-95.8) and AUC of 83.2% (95% CI: 74.5-91.9), reducing workload by 81.3%.
Conclusion:
A machine-assisted screening workflow reduced workload. The workflow extended workload reduction to full-text screening and embedded community researchers throughout, enabling the completion of a scoping review with nuanced eligibility criteria.
Plain Language Summary:
Every year, more scientific papers are published. This makes it hard for researchers to keep up with all the information when they are doing studies that summarize what we know; these studies are called reviews. One of the most time-consuming parts of a review is screening thousands of papers for fit to the review topic. Researchers usually begin with title and abstract screening, which means reading the title and a short summary of each paper (called an abstract) to narrow down, before full-text screening, where they read entire papers to make final inclusion decisions. Artificial intelligence (AI) tools such as machine learning models can help speed up the review process by screening papers. However, many AI tools are developed without knowledge user input, which can improve the relevance of review outcomes and uncover nuances in the review topic, such as assessing the engagement of lay communities in research, which is often inconsistently defined and applied. In this study, we implemented an AI-assisted process to support title and abstract, and full-text screening. We applied the process to a review of approaches used to engage lay communities in modeling studies of infectious diseases. The study was led by knowledge users from community-based organizations in collaboration with academic partners, who together defined the review topic and screened titles and abstracts and full texts. First, we trained a machine learning model on a subset of titles and abstracts that were screened by human reviewers. Then, for full-text screening, we used a two-step approach: we first narrowed down full texts by identifying those that were more likely to fit the review topic based on the words and phrases they used, and then applied a second machine learning model trained on full texts screened by human reviewers. The AI-assisted process helped us screen more than 9,000 papers, reducing screening workload by 48% during title and abstract screening and by 81% during full-text screening. These findings show that combining AI tools with knowledge user input can support reductions in screening workload while accounting for nuanced and inconsistently reported information.