Related Experiment Video
Updated: May 12, 2026

One Dimensional Turing-Like Handshake Test for Motor Intelligence
Published on: December 15, 2010
Benchmarking Human-AI collaboration for common evidence appraisal tools
Tim Woelfle1, Julian Hirt2, Perrine Janiaud3
1Pragmatic Evidence Lab, Research Center for Clinical Neuroimmunology and Neuroscience Basel (RC2NB), Basel, Switzerland; Department of Neurology, University Hospital Basel, Basel, Switzerland; Translational Imaging in Neurology (ThINk), Department of Biomedical Engineering, University Hospital and University of Basel, Basel, Switzerland.
Large language models (LLMs) showed lower accuracy than humans in appraising scientific evidence. Human-AI collaboration, however, improved accuracy and efficiency in evidence appraisal tasks.
Area of Science:
- Artificial Intelligence in Scientific Research
- Medical Evidence Appraisal
- Systematic Review Methodologies
Background:
- Assessing the utility of large language models (LLMs) in evidence appraisal is crucial for optimizing research workflows.
- Current methods for evaluating scientific reporting and methodological rigor are time- and resource-intensive.
Purpose of the Study:
- To quantify the agreement between LLMs and human consensus in appraising systematic reviews and clinical trial designs.
- To identify opportunities for human-AI collaboration to enhance efficiency in evidence appraisal.
Main Methods:
- Five LLMs were used to assess 112 systematic reviews (using PRISMA and AMSTAR criteria) and 56 randomized controlled trials (using PRECIS-2 criteria).
- Agreement was measured between human consensus, individual human raters, individual LLMs, combined LLMs, and human-AI collaboration.
- Ratings were marked as deferred when inconsistencies arose between LLMs or between human raters and LLMs.
Main Results:
- Individual human raters achieved 89% accuracy for PRISMA/AMSTAR and 75% for PRECIS-2.
- Individual LLMs ranged from 38% to 74% accuracy, with combined LLMs showing 64%-89% accuracy but high deferral rates.
- Human-AI collaboration yielded the highest accuracies (80%-96%) with varying deferral rates.
Conclusions:
- LLMs alone performed worse than human appraisers in evaluating scientific evidence.
- Human-AI collaboration shows potential to reduce workload for systematic review reporting and rigor assessments.
- Complex tasks like clinical trial design appraisal (PRECIS-2) were not significantly improved by current LLM-human collaboration.
More Related Videos
Related Concept Videos
Steady Flow of a Fluid Stream
During this process, the momentum of the fluid within the control volume remains constant over the time interval dt. By applying the...
Conservation of Mass in Moving, Nondeforming Control Volume
In the context of a detention basin, the conservation of mass states that the total mass of water entering the basin must equal the mass leaving the basin plus any accumulation of...
Typical Model Studies
Design Example: Creating a Hydraulic Model of a Dam Spillway
Gradually Varying Flow
Rapidly Varying Flow

