Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of large language models in data extraction for evidence synthesis: A systematic review
Ravi Shankar1, Amaevia Lim2, Xu Qian3
1Clinical Research & Innovation Office, Tan Tock Seng Hospital, National Healthcare Group, Singapore; Rehabilitation Research Institute of Singapore, Nanyang Technological University, Singapore.
Journal of Biomedical Informatics
|July 25, 2026
Summary
Large language models (LLMs) show variable but promising accuracy for data extraction in systematic reviews, performing best as assistive tools. Current evidence suggests integrating LLMs into dual-extraction workflows with human oversight for reliable evidence synthesis.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Research
- Evidence Synthesis Methodology
Background:
- Data extraction is a critical yet labor-intensive and error-prone step in systematic reviews.
- Large language models (LLMs) present a potential solution for automating or semi-automating data extraction.
- Comprehensive evaluation of LLM performance in this domain is needed.
Purpose of the Study:
- To systematically review and evaluate the accuracy, reliability, and efficiency of LLMs for data extraction in evidence synthesis.
- To identify optimal strategies for implementing LLMs in systematic review workflows.
- To assess the methodological quality and reporting completeness of studies evaluating LLMs for data extraction.
Main Methods:
- Systematic search of PubMed, Embase, Web of Science, and preprint servers (medRxiv, arXiv) through December 2025.
- Inclusion of studies evaluating LLMs against human reference standards for data extraction, reporting quantitative metrics.
- Independent data extraction and quality assessment by two reviewers using PROBAST+AI and TRIPOD-LLM; narrative synthesis due to heterogeneity.
Main Results:
- Twenty-seven studies evaluated various LLMs (e.g., GPT-4, Claude 3.5, Gemini, Llama), with overall accuracy ranging from 47% to 99.9%.
- Categorical data extraction (74-96%) was more reliable than numerical data (47-88%); Claude 3.5 Sonnet achieved 91.0% accuracy in an assistive role.
- Omissions were the primary error type (60-74%), with low hallucination rates (0.08-6%); reported time savings ranged from 33% to 87%.
Conclusions:
- LLMs demonstrate promising but inconsistent performance for data extraction, supporting their use as assistive tools rather than autonomous extractors.
- Integration into dual-extraction workflows with human verification is recommended for reliable evidence synthesis.
- Future research should focus on standardized benchmarks and prospective comparative studies, particularly for numerical data extraction.
