Related Experiment Video
Updated: Aug 6, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Artificial intelligence in systematic reviews and meta-analyses: Task-specific performance, residual error
Jules Descamps1, Nicolas Bouguennec2, Raphael Porcher3
1Division of Orthopaedic Surgery, Lariboisière hospital, 2 rue Ambroise Paré, 75010, Paris, France; Université Paris Cité, School of Medecine, Paris, France.
Artificial intelligence (AI) assists systematic reviews (SRs) and meta-analyses (MAs) by automating tasks, but human oversight is crucial. AI performance varies by task, necessitating careful benchmarking and quantification of residual errors for reliable evidence synthesis.
Area of Science:
- Medical Informatics
- Evidence Synthesis
- Artificial Intelligence in Research
Background:
- The rapid growth of published research necessitates efficient evidence synthesis methods like systematic reviews (SRs) and meta-analyses (MAs).
- Traditional SR/MA workflows are time-consuming, often exceeding one year.
- Artificial intelligence (AI) offers potential automation but requires rigorous validation.
Purpose of the Study:
- To evaluate AI's reliability across different stages of the SR/MA workflow.
- To define methods for benchmarking AI performance in SR/MA tasks.
- To quantify residual errors after AI-assisted human review.
Main Methods:
- Narrative review of AI applications in SR/MA.
- Analysis of AI performance metrics for specific tasks (searching, screening, data extraction, risk of bias).
- Examination of methods for benchmarking and residual error quantification.
Main Results:
- AI performance is heterogeneous across SR/MA tasks, with significant variability in accuracy.
- Title/abstract screening shows higher maturity (e.g., 99.2% sensitivity), while risk-of-bias assessment is less reliable (kappa 0.51).
- AI significantly reduces time for data extraction but requires refinement to ensure accuracy.
Conclusions:
- AI is best utilized as an assistive technology in hybrid human-machine workflows for SR/MA, not as a fully autonomous tool.
- Task-specific performance evaluation and quantification of residual uncertainty are essential.
- Human oversight remains critical for ensuring the methodological rigor and reliability of evidence syntheses.