Related Experiment Video
Updated: Jan 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Advancing Question-Answering in Ophthalmology With Retrieval-Augmented Generation: Benchmarking Open-Source and
Quang Nguyen1,2,3, Duy-Anh Nguyen4, Khang Dang5
1UCL Institute of Ophthalmology, London, UK.
Retrieval-Augmented Generation (RAG) significantly boosts the accuracy of open-source large language models (LLMs) in ophthalmology question-answering. This approach enhances smaller LLMs, making them suitable for sensitive, resource-limited settings.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Ophthalmology
Background:
- Large language models (LLMs) show promise in medical question-answering.
- Evaluating LLM performance in specialized fields like ophthalmology is crucial.
- Retrieval-Augmented Generation (RAG) combines information retrieval with text generation to improve LLM accuracy.
Purpose of the Study:
- To benchmark open-source and proprietary LLMs in ophthalmology question-answering using RAG.
- To assess the impact of RAG on LLM performance across different models.
- To evaluate the effectiveness of model quantization for efficiency.
Main Methods:
- A dataset of 260 multiple-choice ophthalmology questions from AAO BCSC and OphthoQuestions was used.
- A RAG pipeline with ChromaDB for retrieval and Cohere for reranking was implemented.
- GPT-4-turbo, Llama-3-70B, Gemma-2-27B, and Mixtral-8 × 7B were benchmarked using zero-shot, zero-shot-CoT, and RAG.
- Quantization was applied to open-source models to measure efficiency effects.
Main Results:
- RAG improved GPT-4-turbo accuracy by 10.96-11.54% and open-source models (Llama-3, Gemma-2, Mixtral) by 17.11-23.85%.
- Zero-shot-CoT did not significantly improve model performance.
- 4-bit quantization was as effective as 8-bit while halving resource requirements.
Conclusions:
- RAG significantly enhances LLM accuracy, particularly for smaller open-source models.
- RAG enables privacy-preserving, efficient LLM deployment in resource-constrained environments like hospitals.
- This approach offers a viable alternative to cloud-based LLMs for specialized medical applications.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy