通过检索增强生成在眼科中推进问答:对开源和专有大型语言模型进行基准测试
Quang Nguyen1,2,3, Duy-Anh Nguyen4, Khang Dang5
1UCL Institute of Ophthalmology, London, UK.
Translational vision science & technology
|September 12, 2025
概括
检索增强生成 (RAG) 显著提高了眼科问答中的开源大语言模型 (LLM) 的准确性. 这种方法增强了较小的LLM,使它们适合于敏感的,资源有限的环境.
科学领域:
- 人工智能的人工智能
- 医疗信息学 医疗信息学
- 眼科医生 眼科 眼科
背景情况:
- 大型语言模型 (LLM) 在医学问答方面表现有前途.
- 评估LLM在眼科等专业领域的表现至关重要.
- 检索增强生成 (RAG) 将信息检索与文本生成相结合,以提高LLM的准确性.
研究的目的:
- 用RAG.使用眼科问答来比较开源和专有LLM.
- 通过不同的模型评估RAG对LLM绩效的影响.
- 评估模型量化对效率的有效性.
主要方法:
- 使用了来自AAO BCSC和OphthoQuestions的260个多选眼科问题的数据集.
- 实现了一个RAG管道,使用ChromaDB进行检索和Cohere进行重新排名.
- GPT-4-turbo,Llama-3-70B,Gemma-2-27B和Mixtral-8 × 7B使用零射击,零射击-CoT和RAG进行了基准测试.
- 量化被应用到开源模型来测量效率效应.
主要成果:
- 在RAG中,GPT-4-turbo的精度提高了10.96-11.54%和开源模型 (Llama-3,Gemma-2,Mixtral) 的精度提高了17.11-23.85%.
- 零射击CoT并没有显著改善模型性能.
- 4位量子化与8位一样有效,同时减少了资源需求的一半.
结论:
- RAG显著提高了LLM的准确性,特别是在较小的开源模型中.
- 在医院等资源有限的环境中,RAG可以实现保护隐私,高效的LLM部署.
- 这种方法为专业医疗应用提供了基于云的LLM的可行替代方案.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
3.3K
相关概念视频
Improving Translational Accuracy
3.6K
3.6K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
