Related Experiment Video
Updated: Aug 6, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Techniques, Performance, and Feasibility of Natural Language Processing for Abstract Screening in Evidence Synthesis:
Ravi Shankar1, Ziyu Goh2, Qian Xu3
1Clinical Research & Innovation Office, Tan Tock Seng Hospital, National Healthcare Group, Singapore.
Campbell Systematic Reviews
|July 23, 2026
Summary
Natural language processing (NLP) significantly speeds up systematic review abstract screening, with deep learning models showing high recall and workload reduction. Challenges in data, expertise, and usability remain for full implementation.
Area of Science:
- Information Science
- Computer Science
- Medical Informatics
Background:
- Exponential growth in literature necessitates efficient systematic review screening.
- Manual abstract screening is time-consuming and prone to errors, impacting review timelines.
- Artificial intelligence, including deep learning, offers potential for automating abstract screening.
Purpose of the Study:
- To systematically review natural language processing (NLP) techniques for title and abstract screening in evidence syntheses.
- To assess the performance, including workload reduction and recall, of NLP methods.
- To evaluate the real-world implementation feasibility and identify research gaps.
Main Methods:
- Comprehensive search of multiple databases (PubMed, Web of Science, etc.) and gray literature.
- Inclusion of primary studies on NLP techniques for abstract screening in evidence syntheses.
- Independent screening by two reviewers, data extraction, risk of bias assessment, and narrative synthesis.
Main Results:
- 19 studies met criteria, with most published recently, indicating rapid advancement.
- Deep learning models, particularly BERT variants with transfer learning, showed superior performance (>90% recall, 13-96% workload reduction).
- Implementation barriers include data quality, computational resources, technical expertise, and user interface design.
Conclusions:
- NLP, especially deep learning, shows significant promise for semi-automating abstract screening, offering substantial time savings.
- Overcoming challenges in data, expertise, and user-centered design is crucial for widespread adoption.
- Future work should focus on standardized datasets, prospective evaluations, and user-friendly tool development.
