Related Experiment Video
Updated: Jun 16, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Human-Comparable Sensitivity of Large Language Models in Identifying Eligible Studies Through Title and Abstract
Kentaro Matsui1,2, Tomohiro Utsumi2,3, Yumi Aoki4
1Department of Clinical Laboratory, National Center Hospital, National Center of Neurology and Psychiatry, Kodaira, Japan.
This study introduces a 3-layer screening method using GPT-4 for systematic reviews, significantly improving efficiency and maintaining high sensitivity for relevant records. The AI-powered approach streamlines the screening process, making systematic reviews more accessible.
Area of Science:
- Bibliometrics and Information Science
- Artificial Intelligence in Research
- Medical Informatics
Background:
- Systematic review screening is time-consuming and labor-intensive.
- Previous AI solutions reduced workload but risked excluding relevant studies.
- Developing efficient and accurate screening methods is crucial for evidence synthesis.
Purpose of the Study:
- To evaluate a 3-layer screening method using GPT-3.5 and GPT-4 for systematic reviews.
- To streamline title and abstract screening for systematic reviews.
- To maximize sensitivity in identifying relevant records during systematic review screening.
Main Methods:
- Applied a 3-layer screening process (research design, patients, interventions) using GPT-3.5 and GPT-4.
- Screened 1381 and 3146 records from two systematic reviews on bipolar disorder treatment.
- Utilized tailored prompts and an automated GPT-4 flow for information extraction and screening optimization.
Main Results:
- GPT-3.5 and GPT-4 screened approximately 110 records per minute.
- GPT-4 achieved high sensitivities and specificities (e.g., 0.962/0.996 in study 1, 0.943/0.855 in study 2) after exclusions.
- AI screening sensitivities aligned with human evaluators, with GPT-4 showing fewer domain-knowledge-related errors than GPT-3.5.
Conclusions:
- The 3-layer screening method using GPT-4 demonstrates practical feasibility for systematic reviews.
- The approach shows acceptable sensitivity and specificity for efficient screening.
- Further research is needed to generalize the method across diverse settings and establish operational feasibility.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
09:35A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017