Related Experiment Video
Updated: Apr 10, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
"Can a chatbot be used in the full-text screening in a systematic review?"
André Miguel Martins1, Luis Félix Valero Juan2, Adriana Oliveira3
1LAQV/REQUIMTE, Escola Superior de Saúde, Instituto Politécnico do Porto, Rua Dr. António Bernardino de Almeida, 4200-072 Porto, Portugal.
Introduction:
Large language model-based artificial intelligence tools are increasingly explored to support systematic reviews, yet evidence regarding their reliability in full-text screening remains limited. This study evaluated the performance of two versions of ChatGPT (4.0 and 5.0) compared with human reviewers during article selection for a systematic review on influenza vaccine effectiveness.
Methods:
A total of 170 full-text articles were independently assessed for eligibility using predefined inclusion and exclusion criteria. Human reviewers served as the gold standard. ChatGPT 4.0 and 5.0 were prompted using standardized instructions mirroring the review protocol. Agreement with human decisions was evaluated using accuracy, sensitivity, specificity, precision, F1-score, and Cohen's κ. Intra-model reproducibility was assessed for ChatGPT 5.0.
Results:
ChatGPT 4.0 achieved an accuracy of 0.71 (95% CI: 0.64-0.78) and a Cohen's κ of 0.43, indicating moderate agreement with human reviewers. ChatGPT 5.0 demonstrated improved performance, with accuracy increasing 0.06 to 0.77 (95% CI: 0.70-0.83), sensitivity of 0.87, specificity of 0.70, and κ of 0.55, corresponding to moderate-to-substantial agreement. Intra-model reproducibility for ChatGPT 5.0 showed 80% agreement (κ = 0.60), indicating partial but imperfect consistency.
Conclusions:
ChatGPT 5.0 outperformed ChatGPT 4.0 in full-text screening accuracy and reproducibility, approaching but not matching human performance. These findings support the use of current LLMs as decision-support tools rather than autonomous reviewers in systematic reviews. Transparent reporting of model versions, prompts, and input quality is essential to ensure credible AI-assisted evidence synthesis.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Non-equilibrium in the Cell
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...