Related Experiment Video
Updated: May 9, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
461
Use of Retrieval-Augmented Large Language Model for COVID-19 Fact-Checking: Development and Usability Study
Hai Li1, Jingyi Huang1, Mengmeng Ji2
1School of Economics and Management, Shanghai University of Sport, Shanghai, China.
Journal of Medical Internet Research
|April 30, 2025
Summary
Retrieval-augmented generation (RAG) with large language models (LLMs) significantly improves COVID-19 fact-checking accuracy. Advanced RAG models like CRAG and SRAG drastically reduce misinformation and hallucinations for reliable information verification.
Area of Science:
- Artificial Intelligence
- Public Health Informatics
- Computational Linguistics
Background:
- The COVID-19 pandemic fueled an "infodemic" of misinformation, overwhelming traditional fact-checking.
- Large language models (LLMs) offer scalable solutions but are prone to generating hallucinations.
- Limitations in LLM reliability hinder effective, large-scale misinformation combat.
Purpose of the Study:
- To enhance COVID-19 fact-checking accuracy and reliability.
- To address LLM hallucination and context inaccuracy using retrieval-augmented generation (RAG).
- To evaluate RAG-enhanced LLM performance against misinformation.
Main Methods:
- Developed RAG-enhanced models (naïve RAG, LOTR-RAG, CRAG, SRAG) integrated with GPT-4.
- Utilized a dataset of ~130,000 COVID-19 peer-reviewed papers for context.
- Evaluated models on real-world and synthesized datasets (500 claims each) for accuracy, F1-score, precision, and sensitivity.
Main Results:
- RAG models significantly improved accuracy over baseline GPT-4 on both datasets.
- CRAG and SRAG models achieved the highest accuracies (0.972 and 0.973 on real-world; 0.978 on synthesized).
- RAG consistently reduced hallucinations and improved contextual accuracy, especially CRAG and SRAG.
Conclusions:
- Integrating RAG systems with LLMs substantially boosts automated fact-checking accuracy and relevance.
- This approach offers rapid, reliable information verification to combat public health misinformation.
- RAG enhances transparency by citing sources, crucial for trust during health crises.
Related Concept Videos
Leaky Scanning
5.0K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.0K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Single Nucleotide Polymorphisms-SNPs
13.6K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
13.6K

