用检索增强的大型语言模型进行COVID-19事实检查:开发和可用性研究
Hai Li1, Jingyi Huang1, Mengmeng Ji2
1School of Economics and Management, Shanghai University of Sport, Shanghai, China.
Journal of medical Internet research
|April 30, 2025
概括
采用大型语言模型 (LLM) 的检索增强生成 (RAG) 显著提高了COVID-19事实检查的准确性. 像CRAG和SRAG这样的先进RAG模型大大减少了错误信息和幻觉,以便可靠地验证信息.
科学领域:
- 人工智能的人工智能
- 公共卫生信息学 公共卫生信息学
- 计算语言学 计算语言学
背景情况:
- 随着COVID-19的流行,错误信息的"传染病"蔓延,压倒了传统的事实核查.
- 大型语言模型 (LLM) 提供可扩展的解决方案,但容易产生幻觉.
- 在LLM可靠性的局限性阻碍有效,大规模的错误信息打击.
研究的目的:
- 为了提高COVID-19事实核查的准确性和可靠性.
- 为了解决LLM幻觉和上下文不准确性,使用检索增强生成 (RAG).
- 为了评估RAG增强的LLM绩效与错误信息对抗.
主要方法:
- 开发了RAG增强模型 (原始RAG,LOTR-RAG,CRAG,SRAG) 与GPT-4集成.
- 利用了~13万篇COVID-19同行评审论文的数据集,以获得上下文.
- 在真实世界和合成数据集 (每个500个索赔) 上评估模型的准确性,F1分数,精度和灵敏度.
主要成果:
- 在两个数据集上,RAG模型显著提高了比基线GPT-4的准确性.
- CRAG和SRAG模型实现了最高的精度 (0.972和0.973在现实世界;0.978在合成).
- RAG 持续减少了幻觉和提高了语境准确性,特别是 CRAG 和 SRAG.
结论:
- 将RAG系统与LLM集成,大大提高了自动事实检查的准确性和相关性.
- 这种方法提供了快速,可靠的信息验证,以打击公共卫生错误信息.
- 通过引用来源,RAG提高了透明度,这对于健康危机期间的信任至关重要.
相关概念视频
Leaky Scanning
5.0K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.0K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Single Nucleotide Polymorphisms-SNPs
13.6K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
13.6K


