在RAG系统性能评估和ChatGPT优化中的医疗QA对话数据集.
Muretijiang Muhetaer1, Ailimulati Yusupu2, Wang Yifan2
1School of Information Management, Wuhan University, Wuhan, 430072, China. 2019281040215@whu.edu.cn.
Scientific reports
|December 24, 2025
概括
中国的医生与患者的对话,通过Retrieval-Augmented Generation (RAG) 增强了临床问题答案. 对话数据显著改善了结果,优化的检索策略为可靠的医疗问题答案系统提供了最佳平衡.
科学领域:
- 人工智能的人工智能
- 自然语言处理自然语言处理.
- 医疗信息学 医疗信息学
背景情况:
- 临床问答 (QA) 系统旨在提供准确的医疗信息.
- 检索增强生成 (RAG) 通过结合外部知识来增强大型语言模型.
- 优化检索源对于改善RAG性能在医学等专业领域至关重要.
研究的目的:
- 评估中国医生与患者对话作为临床QA中RAG检索来源的有效性.
- 为了比较医疗QA的各种检索策略和高级语言模型 (GPT-4o,GPT-5) .
- 确定影响医学领域RAG表现的关键因素.
主要方法:
- 利用中国医生与病人对话作为RAG的检索库.
- 实施和比较的检索方法:密集检索,交叉编码器重新排名,相互排名融合 (RRF) 和级联RRF→Rerank.
- 通过使用自动指标 (ROUGE,BERTScore) 和不同语言模型 (ChatGPT-3.5,GPT-4o,GPT-5) 的专家人类评估来评估性能.
主要成果:
- 基于对话的检索显著提高了生成质量,而不是直接提示 (ROUGE-1-f: +12.6%,BERTScore_F1: +1.5%).
- 只有Rerank的策略提供了最佳的准确度-延迟平衡;级联管道没有提供额外的好处.
- GPT-4o表现出卓越的自动指标和较低的延迟,而GPT-5获得了略高的人类偏好分数.
结论:
- 中国的医生与患者对话是改善临床质量保证中的RAG的有效检索来源.
- 数据表示和元数据结构对RAG性能比检索算法复杂性更为关键.
- 结果为使用RAG.使用可靠的医疗质量保证系统的部署提供了实际指导.
相关概念视频
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K

