Related Experiment Video
Updated: May 6, 2026

Assessment and Communication for People with Disorders of Consciousness
Published on: August 1, 2017
Reference Hallucination Score for Medical Artificial Intelligence Chatbots: Development and Usability Study
Fadi Aljamaan1, Mohamad-Hani Temsah1, Ibraheem Altamimi1
1College of Medicine, King Saud University, Riyadh, Saudi Arabia.
A new reference hallucination score (RHS) evaluates AI chatbot citations in medical research. Elicit and SciSpace showed low hallucination, while ChatGPT and Bing exhibited critical levels, highlighting the need for verification tools.
Area of Science:
- Medical research
- Artificial intelligence
- Scientific writing
Background:
- AI chatbots are increasingly used by healthcare practitioners.
- AI chatbot outputs, including references, may contain hallucinations.
- These inaccuracies raise concerns about AI reliability in medical contexts.
Purpose of the Study:
- To introduce a novel reference hallucination score (RHS) for assessing the authenticity of AI chatbot-generated citations.
- To quantify and compare citation hallucination across different AI chatbots.
Main Methods:
- Six AI chatbots were tested with 10 medical prompts each, requesting 10 references per prompt.
- The RHS was developed, considering bibliographic details and relevance to prompt keywords.
- RHS was calculated for individual references, prompts, and prompt types (basic vs. complex).
Main Results:
- Bard did not generate references. Elicit and SciSpace had the lowest RHS (1), while ChatGPT 3.5 and Bing had the highest (11).
- Reference relevancy to prompt keywords was the most common hallucination (61.6%).
- AI chatbots produced significantly higher hallucination scores for complex or scenario-based prompts.
Conclusions:
- The study highlights significant variations in AI chatbot citation authenticity, necessitating robust evaluation tools.
- Elicit and SciSpace demonstrated minimal hallucination, whereas ChatGPT and Bing showed critical levels.
- The proposed RHS can aid in improving the reliability and trustworthiness of AI in medical research.
More Related Videos
06:02Evaluating Usability Aspects of a Mixed Reality Solution for Immersive Analytics in Industry 4.0 Scenarios
Published on: October 6, 2020
08:36The Immersive Cleveland Clinic Virtual Reality Shopping Platform for the Assessment of Instrumental Activities of Daily Living
Published on: July 28, 2022