Related Experiment Video
Updated: Apr 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Fact-Checking Large Language Model Responses to a Health Care Prompt: Comparative Study
Padhraig Ryan1, Orla Davoren, Glyn Elwyn2
1Pharmaceutical Society of Ireland, Dublin, Ireland.
Large language models (LLMs) show promise in healthcare fact-checking and patient interaction. While efficient, human oversight is crucial for ensuring complete accuracy and patient safety.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing
- Clinical Decision Support
Background:
- Large language models (LLMs) offer potential applications in healthcare, including patient education and diagnosis.
- Current evaluations of LLMs in healthcare settings are limited.
- This study addresses the need for rigorous assessment of LLM capabilities in clinical contexts.
Purpose of the Study:
- To evaluate the accuracy and efficiency of automated fact-checking using two distinct LLMs.
- To demonstrate how an LLM can assist patients in refining prompts for improved clinical safety.
- To compare LLM performance against human expert evaluations.
Main Methods:
- A comparative study involving two LLMs (GPT-4o, OpenBioLLM-70B) and three human experts.
- A clinical scenario focused on retinoid safety for acne treatment, involving prompt refinement and fact-checking.
- Evaluation of LLM fact-checking accuracy and time efficiency on 20 diverse clinical statements.
- Outcome measures included accuracy percentage, time to fact-check, and prompt redrafting success.
Main Results:
- GPT-4o and OpenBioLLM-70B achieved 86% agreement with experts in the clinical scenario, though with some omissions regarding critical safety information (isotretinoin, folic acid).
- For 20 clinical statements, GPT-4o demonstrated 100% accuracy compared to human experts, while OpenBioLLM-70B achieved 95% accuracy.
- LLMs significantly reduced fact-checking time compared to human experts (seconds/minutes vs. 18 minutes on average).
Conclusions:
- GPT-4o can effectively enhance patient prompts for better health information capture and safety.
- Both evaluated LLMs provide efficient fact-checking with accuracy levels approaching those of human experts.
- Human expert review remains essential to verify accuracy and ensure comprehensive patient safety, especially for critical information.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Healthcare Agencies II
Parish nursing is a growing specialty nursing profession that focuses on holistic healthcare, health promotion, and illness prevention. It blends professional nursing practice with a health ministry, focusing on health and healing within the context of a Christian community. Parish nurses serve as health educators, referral sources,...
Healthcare Agencies I
Health Literacy
Preventive Healthcare Services
Models of Health Promotion and Illness Prevention II
The agent-host-environment model states that disease results...
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...