Related Experiment Video
Updated: Jan 12, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Human-led and artificial intelligence-automated critical appraisal of systematic reviews: Comparative evaluation
Lucija Gosak1, Gregor Štiglic2, Wilson Wai San Tam3
1Faculty of Health Sciences, University of Maribor, Maribor 2000, Slovenia.
Aim:
To evaluate and compare human-led and artificial intelligence-automated critical appraisal of evidence.
Background:
Critical appraisal is essential in evidence-based practice, yet many nurses lack the skills to perform it. Large language models offer potential support, but their role in critical appraisal remains underexplored.
Design:
We conducted a comparative study to evaluate the performance of five commonly used large language models versus two human reviewers in appraising four systematic reviews on interventions to reduce medication administration errors.
Methods:
We compared large language models and two human reviewers in independently appraising four systematic reviews using the JBI Critical Appraisal Checklist. These models were Perplexity Sonar (Pro), Claude 3.7 Sonnet, Gemini 2.0 Flash, GPT-4.5 and Grok-2. All models received identical full texts and standardized prompts. Responses were analyzed descriptively and agreement was assessed using Cohen's Kappa.
Results:
Large language models showed full agreement with human reviewers on five of 11 JBI items. Most disagreements occurred in appraising search strategy, inclusion criteria and publication bias. The agreement between human reviewers and large language models ranged from slight to moderate. The highest level of agreement was observed with Claude (κ = 0.732), while the lowest level was observed with Gemini (κ = 0.394).
Conclusion:
Large language models can support aspects of critical appraisal evidence but lack contextual reasoning and methodological insight required for complex judgments. While Claude 3.7 Sonnet aligned most closely with human reviewers, human oversight remains essential. Large language models should serve as adjuncts and not substitutes for evidence-based practice.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:08Estimate the Cognitive Load Using Electrocardiographic Measure: A Human-AI Collaborative Task
Published on: December 5, 2025
Related Concept Videos
Humanistic Psychology
This approach...
Introduction to Cognitive Psychology
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
Non-equilibrium in the Cell
Reason and Intuition