Related Experiment Video
Updated: Jun 23, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of two large language models for data extraction in evidence synthesis
Amanda Konet1, Ian Thomas1, Gerald Gartlehner1,2
1Social, Statistical, and Environmental Sciences, RTI International, Research Triangle Park, North Carolina, USA.
Large language models (LLMs) show promise for evidence synthesis data extraction. Claude 2 demonstrated higher accuracy than GPT-4, largely due to GPT-4
Area of Science:
- Artificial Intelligence in Scientific Research
- Biomedical Informatics
- Evidence Synthesis Methodologies
Background:
- Accurate data extraction is crucial for reliable evidence synthesis.
- Large language models (LLMs) present potential for automating data extraction.
- Uncertainty exists regarding the optimal LLM for evidence synthesis tasks.
Purpose of the Study:
- To compare the performance of two widely available LLMs, Claude 2 and GPT-4, for data extraction in evidence synthesis.
- To evaluate the accuracy of LLM-driven data extraction from full-text scientific articles.
Main Methods:
- Two LLMs (Claude 2, GPT-4) were used to extract pre-specified data elements from 10 published articles.
- Full study PDFs were processed using browser versions of the LLMs.
- GPT-4 required a third-party plugin for PDF parsing, while Claude 2 did not.
- Accuracy was assessed by comparing LLM outputs to previously extracted data.
Main Results:
- Claude 2 achieved high accuracy (96.3%) in data extraction.
- GPT-4, with its PDF parsing plugin, showed lower accuracy (68.8%), with most errors attributed to the plugin.
- Both LLMs accurately identified missing data elements and handled un-reported information.
- When provided with selected text, Claude 2 and GPT-4 achieved 98.7% and 100% accuracy, respectively.
Conclusions:
- LLMs have significant potential to enhance data extraction efficiency in evidence synthesis.
- Accurate PDF parsing is a critical factor influencing LLM performance.
- Human oversight remains essential for validating LLM-generated extractions.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Extraction: Advanced Methods
Improving Translational Accuracy
Evolutionary Relationships through Genome Comparisons
Goodness-of-Fit Test
Crossover Experiments
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...