Related Experiment Video
Updated: May 26, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated full-text screening and accelerated reviews using large language models with context-aware agents: an
Yongxin Ye1, Martina Colombo2, Jennifer Meessen2
1NOVEL Translational Medicine, Novo Nordisk A/S, Vandtårnsvej 108, Søborg 2860, Denmark.
European Heart Journal. Digital Health
|May 25, 2026
Summary
Artificial intelligence (AI) tools using large language models (LLMs) enhance scientific literature reviews. Our AI tool automates full-text screening, improving accuracy and reducing manual effort for biomarker identification in heart failure.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Medicine
- Cardiology Research
Background:
- Scientific literature reviews are crucial for identifying relevant patient populations and biomarkers.
- Manual screening of full-text publications is labor-intensive and prone to human error.
- Large language models (LLMs) offer potential for automating and improving the efficiency of literature reviews.
Purpose of the Study:
- To develop and evaluate an AI-based tool for automating full-text screening in scientific literature reviews.
- To improve the accuracy and efficiency of identifying relevant publications based on complex criteria, specifically focusing on biomarkers in heart failure with reduced ejection fraction (HFrEF).
Main Methods:
- A literature review was conducted using the Population, Intervention-biomarkers, Comparison, Outcome (PICOT) framework.
- An AI tool was developed using retrieval-augmented generation (RAG) and agent-based methods to process 5405 publications.
- LLaMA 3.3 70B was selected for its performance, and ground truth standards were established to evaluate the AI tool against human reviewers.
Main Results:
- The AI tool demonstrated high accuracy, precision, and recall, with LLaMA 3.3 70B achieving 82% accuracy, 71% precision, and 100% recall in initial screenings.
- Validation results showed a sensitivity of 91.4% and a specificity of 53.2%.
- The AI tool outperformed human reviewers in F1 score and interrater reliability, achieving 100% consistency across multiple runs.
Conclusions:
- AI tools, particularly LLMs, can significantly reduce labor-intensive efforts in scientific literature reviews while maintaining accuracy.
- The developed AI tool demonstrates superior inter-rater agreement compared to human reviewers.
- Automated full-text screening using AI holds promise for accelerating research discovery in fields like cardiology.
