Related Experiment Video
Updated: Mar 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Large language models show promising performance for some systematic review tasks but call for cautious
Florian Laignelot1, Guillaume L Martin1, Mohamad Ossman1
1Sorbonne Université, INSERM, Institut Pierre Louis d'Epidémiologie et de Santé Publique, UMR-S 1136, AP-HP, Hôpital Pitié-Salpêtrière, Département de Santé Publique, Paris, France.
Journal of Clinical Epidemiology
|March 14, 2026
Summary
Large language models (LLMs) show promise for automating systematic review tasks like screening, with newer models performing better. Careful implementation is key for integrating LLMs into systematic reviews.
Area of Science:
- Biomedical informatics
- Artificial intelligence in healthcare
- Systematic review methodology
Background:
- The increasing volume of biomedical literature poses challenges for conducting systematic reviews.
- Automation of systematic review processes is needed to improve efficiency.
Purpose of the Study:
- To evaluate the performance of large language models (LLMs) in automating systematic review and meta-analysis steps.
- To assess LLM capabilities across various stages of the systematic review workflow.
Main Methods:
- A systematic review of studies assessing LLM performance in systematic reviews.
- Searched PubMed, Embase, Cochrane Library, and preprint platforms up to January 14, 2025.
- Extracted data and assessed risk of bias, analyzing performance metrics like positive (PPA) and negative percent agreement (NPA).
Main Results:
- Included 63 studies with 148 LLM performance assessments, primarily on GPT models.
- LLMs demonstrated high performance in Title/Abstract (PPA 0.92, NPA 0.89) and Full-Text screening (PPA 0.93, NPA 0.92).
- Data extraction (median accuracy 0.95) and Risk of Bias assessment (median accuracy 0.62) showed variable but promising results, with newer LLMs outperforming older ones.
Conclusions:
- Large language models show significant potential for automating repetitive tasks in systematic reviews, especially screening.
- Integration of LLMs requires careful implementation and appropriate safeguards for reliable use in evidence synthesis.
Keywords:
Artificial intelligenceLarge language modelsMeta-analysesMethodologyScreeningSystematic reviewsMore Related Videos
Related Concept Videos
Improving Translational Accuracy
15.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.4K
Improving Translational Accuracy
3.7K
3.7K

