Related Experiment Video
Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Data Extractions Using a Large Language Model (Elicit) and Human Reviewers in Randomized Controlled Trials: A
Joleen Bianchi1,2, Julian Hirt1,3,4, Magdalena Vogt1
1Department of Health Eastern Switzerland University of Applied Sciences St. Gallen Switzerland.
Elicit, an AI tool, partially extracts data for systematic reviews but requires human verification for accuracy. Human reviewers are essential to ensure complete and correct data extraction from randomized controlled trials.
Area of Science:
- Medical research methodology
- Artificial intelligence in healthcare
Background:
- Systematic reviews are crucial for evidence-based medicine but are labor-intensive.
- Artificial intelligence (AI) tools like Elicit may streamline systematic review processes, particularly data extraction.
- The accuracy and performance of Elicit for data extraction have not been independently validated.
Purpose of the Study:
- To compare the accuracy of data extraction from randomized controlled trials (RCTs) using the AI tool Elicit versus human reviewers.
- To assess Elicit's performance across various data variables within RCTs.
Main Methods:
- A comparative analysis was conducted on 20 RCTs from diverse healthcare topics and 11 countries.
- Data extraction was performed by both Elicit and human reviewers for predefined variables: study objectives, sample characteristics/size, study design, interventions, outcomes, and intervention effects.
- Extracted data were classified as "more," "equal to," "partially equal," or "deviating" compared to human extraction, using the STROBE checklist for reporting.
Main Results:
- Elicit demonstrated partial accuracy, with "more" data extracted in 29.3%, "equal" in 20.7%, "partially equal" in 45.7%, and "deviating" in 4.3% of cases across seven variables.
- Elicit excelled in extracting study design (100% "more") and sample characteristics (45% "more").
- For complex variables like "intervention effects" and "interventions," Elicit's extractions were less detailed, with 95% rated as "partially equal."
Conclusions:
- Elicit can partially extract data for systematic reviews but requires human oversight for nuanced variables.
- Human reviewers remain essential for ensuring the completeness and accuracy of data extraction, especially for intervention details and effects.
- While AI tools can facilitate data extraction, verification by human reviewers is necessary for reliable systematic reviews.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Group Design
Bioequivalence Experimental Study Designs: Repeated Measures, Cross-Over, Carry-Over, and Latin Square Designs
Bioequivalence Experimental Study Designs: Completely Randomized and Randomized Block Designs
Data Collection by Experiments
An example of the experimental method is a public...
Randomized Experiments
Simple randomization
Simple...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...