Related Experiment Video
Updated: Jun 25, 2025

Assessment and Communication for People with Disorders of Consciousness
Published on: August 1, 2017
Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis
Mikaël Chelli1, Jules Descamps2, Vincent Lavoué1
1Institute for Sports and Reconstructive Bone and Joint Surgery, Groupe Kantys, Nice, France.
Large language models (LLMs) show poor performance in generating accurate references for systematic reviews, with high hallucination rates. Researchers must validate all LLM-generated citations to ensure reliability in academic writing.
Area of Science:
- Artificial Intelligence in Scientific Research
- Bibliometrics and Information Science
- Medical Literature Analysis
Background:
- Large language models (LLMs) offer potential for automating literature searches in systematic reviews.
- Concerns exist regarding LLM reliability and the persistent generation of unsupported (hallucinated) content.
- Assessing LLM performance in scientific reference generation is crucial for academic integrity.
Purpose of the Study:
- To evaluate the reference-generating capabilities of LLMs, specifically ChatGPT and Bard (now Gemini).
- To compare LLM performance against human-conducted systematic reviews in scientific writing.
- To quantify accuracy, recall, precision, and hallucination rates of LLM-generated references.
Main Methods:
- LLMs (GPT-3.5, GPT-4, Bard) were tested using inclusion criteria from systematic reviews on shoulder rotator cuff pathology.
- Performance was measured against gold-standard references from human-conducted systematic reviews.
- Key metrics included precision, recall, F1-score, and hallucination rate, with hallucinations defined by incorrect title, author, or year.
Main Results:
- Bard exhibited 0% precision and a 91.4% hallucination rate; GPT-3.5 and GPT-4 showed low precision (9.4%-13.4%) and high hallucination rates (28.6%-39.6%).
- Recall rates for GPT-3.5 and GPT-4 were also low (11.9%-13.7%), with Bard failing to retrieve relevant papers.
- LLMs demonstrated biases in study type, participant criteria, and geographical/open-access representation, alongside significant hallucination issues.
Conclusions:
- Current LLMs are not recommended as primary tools for systematic reviews due to poor accuracy and high hallucination rates.
- All references generated by LLMs require rigorous validation by human researchers.
- Improvements in LLM training and functionality are necessary before their reliable use in academic research.
More Related Videos
06:15Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
06:19Comparison of Three Clinical Stereoscopic Methods for Measuring Binocular Visual Function During Amblyopic Treatment in Unilateral Amblyopia
Published on: September 27, 2024
Related Concept Videos
Hallucinogens and Psychedelics
Marijuana, derived from the dried leaves and flowers of the hemp plant, contains...
Positive Symptoms Schizophrenia: Hallucinations and Delusions
Hallucinations
Hallucinations in...
Blind Procedures