Related Experiment Video
Updated: Jul 11, 2026

Visualization of Intensity Levels to Reduce the Gap Between Self-Reported and Directly Measured Physical Activity
Published on: March 7, 2019
Artificial intelligence versus human consensus: A concordance analysis in the screening of studies for evidence
Sebastián Rodríguez1, Catalina León-Prieto2, María Fernanda Rodríguez-Jaime3
1Facultad de Ciencias del movimiento, Programa de Fisioterapia, Universidad FUCS, Bogotá, Colombia.
Objective:
To assess the agreement between the ChatGPT Plus (GPT 4.1) version and human consensus during the screening of studies in four different evidence synthesis projects.
Methods:
A comparative design was used to analyze the degree of agreement between ChatGPT Plus (GPT 4.1) and human reviewers in the study selection process for two systematic reviews with meta-analyses, one scoping review, and one literature review. Human screening was performed independently using the Rayyan platform, while the artificial intelligence was provided with predefined eligibility criteria and protocols. Screening decisions were compared using Cohen's kappa coefficient, sensitivity, and specificity, using Stata 18.
Results:
In the systematic reviews with meta-analyses (SR1 and SR2), agreement was high (κ = 0.73 and 0.86), with sensitivity ≥0.88 and specificity ≥0.99, indicating high reliability in excluding irrelevant studies. In contrast, in the scoping review (SR3) and the literature review (NR1), agreement was moderate (κ = 0.56 and 0.59), with lower positive predictive values (≤0.52), suggesting a higher risk of overdetection.
Conclusion:
Overall, the results suggest that AI-assisted screening may serve as a reliable support tool or triage aid in reviews with well-defined inclusion criteria, rather than fully replacing manual screening. However, in more exploratory or interpretive contexts, human oversight remains necessary to ensure the accuracy of the selection process.

