Related Experiment Video
Updated: Jun 20, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Development and Evaluation of a Retrieval-Augmented Large Language Model Framework for Ophthalmology
Ming-Jie Luo1, Jianyu Pang1, Shaowei Bi1
1State Key Laboratory of Ophthalmology, Zhongshan Ophthalmic Center, Sun Yat-sen University, Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science, Guangdong Provincial Clinical Research Center for Ocular Diseases, Guangzhou, China.
Augmenting large language models (LLMs) with ophthalmic knowledge bases significantly improved their accuracy and utility for clinical questions. This retrieval-augmented LLM offers a safe and practical tool for healthcare professionals.
Area of Science:
- Artificial Intelligence in Medicine
- Ophthalmology
- Natural Language Processing
Background:
- Augmenting large language models (LLMs) with knowledge bases can enhance medical domain performance.
- Practical, privacy-conscious local implementations of LLMs are needed for healthcare professionals.
- Existing LLMs may not meet the specific needs of specialized medical fields like ophthalmology.
Purpose of the Study:
- To develop an accurate and cost-effective local implementation of a retrieval-augmented LLM.
- To mitigate privacy concerns associated with LLM deployment in healthcare settings.
- To enhance the accessibility and practical utility of LLMs for healthcare professionals.
Main Methods:
- Developed ChatZOC, a retrieval-augmented LLM framework, by enhancing a baseline LLM with a comprehensive ophthalmic dataset (over 30,000 knowledge pieces).
- Benchmarked ChatZOC against 10 representative LLMs, including GPT-4 and GPT-3.5 Turbo, using 300 clinical ophthalmology questions.
- Evaluated LLM performance on accuracy, utility, and safety using a double-masked approach with medical experts and researchers.
Main Results:
- The retrieval-augmented LLM achieved a human ranking score of 0.60, significantly improving upon the baseline model (0.48) and comparable to GPT-4 (0.61).
- Scientific consensus accuracy for the retrieval-augmented LLM was 84.0%, a substantial increase from the baseline (46.5%) and comparable to GPT-4 (79.2%).
- The study demonstrated statistically significant improvements in performance metrics due to knowledge base augmentation.
Conclusions:
- Integrating high-quality knowledge bases significantly improves LLM performance in specialized medical domains.
- Augmented LLMs, like ChatZOC, hold transformative potential for clinical practice by providing reliable and safe information.
- Further research is warranted to explore the real-world application and broader integration of such augmented LLM frameworks in healthcare.
More Related Videos
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022