Related Experiment Video
Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
AI-assisted literature screening: A hybrid approach using large language models and retrieval-augmented generation
Yiming Li1, Xinsong Du1, Yifei Wang2
1Department of Medicine, Harvard Medical School, Boston, MA 02115, United States; Division of General Internal Medicine and Primary Care, Department of Medicine, Brigham and Women's Hospital, Boston, MA 02115, United States.
This study introduces a large language model (LLM) approach for automated literature screening, significantly improving efficiency and accuracy. The hybrid method combines retrieval-augmented generation (RAG) and ensemble strategies for enhanced biomedical evidence synthesis.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Healthcare
- Medical Literature Analysis
Background:
- Manual systematic literature review is time-consuming and difficult to scale.
- Large Language Models (LLMs) offer potential for automating literature screening.
- Developing efficient and accurate LLM-based methods is crucial for evidence synthesis.
Purpose of the Study:
- To enhance the efficiency and accuracy of literature screening using LLM-based approaches.
- To explore rule-based preprocessing, retrieval-augmented generation (RAG), and ensemble strategies for LLM-based screening.
- To evaluate the performance of different LLMs and prompting techniques in identifying relevant biomedical literature.
Main Methods:
- A hybrid framework combining RAG prompting with LLM classification was developed.
- Evaluated DeepSeek-V3, Deepseek-R1, and GPT-4o on 6331 biomedical articles using binary, RAG, and justification-based prompting.
- Prioritized recall for maximizing relevant study inclusion, alongside precision, specificity, and NPV.
- Tested generalizability on ten additional topics related to cancer immunotherapy and LLMs in medicine.
Main Results:
- The hybrid approach (rule-based preprocessing + DeepSeek R1 + RAG prompting) achieved a recall of 0.77 and NPV of 1.00.
- Ensemble methods achieved perfect precision (1.00) but comparable performance in other metrics, except F1.
- DeepSeek-R1 demonstrated strong generalizability with the highest F1-score (0.93) and accuracy (0.88) on additional topics.
Conclusions:
- An LLM-based approach integrating RAG prompting and ensemble strategies significantly enhances literature screening accuracy and scalability.
- This method provides a foundation for advancing LLM-driven evidence synthesis in biomedical research.
- The findings support improved clinical decision support through more efficient literature review.
Related Concept Videos
Non-equilibrium in the Cell
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Improving Translational Accuracy
