Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Performance of Large Language Models in Answering Healthcare Delivery Questions: A Quantitative Cross-Sectional

Mohsen Khosravi1, Zahra Zamaninasab2, Fatemeh Khosravi3

  • 1Social Determinants of Health Research Center Birjand University of Medical Sciences Birjand Iran.

Health Science Reports
|June 12, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Systematic Review on the Outcomes of Implementing Public-Private Partnerships in Healthcare Systems: The Experience of the Middle East.

Health science reports·2026
Same author

Treatment of Nausea and Vomiting With Zingiber officinale in Pregnancy: An Umbrella Review of Systematic Reviews With Meta-Synthesis.

Phytotherapy research : PTR·2026
Same author

A Systematic Review of Factors Affecting Utilization of Decision Support Systems: The Interplay Between Technology, Users, and the Healthcare Environment.

Health science reports·2026
Same author

Effect of virtual reality (VR) technology on anxiety control and acrophobia reduction: A randomized controlled trial in Iran.

Global mental health (Cambridge, England)·2026
Same author

Investigation of the Role of Political Participation in Health: A Scoping Review.

Health science reports·2026
Same author

A systematic review of the limitations of large language models in generating healthcare content.

PLOS digital health·2026

Large Language Models (LLMs) show promise in answering healthcare delivery questions. Gemini 2.5 demonstrated the best performance in a recent study evaluating AI chatbots for healthcare queries.

Area of Science:

  • Artificial Intelligence
  • Healthcare Informatics
  • Natural Language Processing

Background:

  • Large Language Models (LLMs) are increasingly utilized across diverse fields, showing positive outcomes.
  • Understanding health service delivery is crucial for optimizing healthcare systems.
  • This study investigates the efficacy of LLM-based chatbots in addressing healthcare delivery queries.

Purpose of the Study:

  • To analyze and compare the performance of leading LLM-based chatbots in answering healthcare delivery questions.
  • To evaluate the accuracy and reliability of AI models in a healthcare context.
  • To identify the most effective AI chatbot for healthcare-related inquiries.

Main Methods:

  • A validated questionnaire was administered to a sample of LLM-based chatbots, including GPT-4.1-mini, Gemini 2.5, Copilot 2025, and Perplexity.
Keywords:
artificial intelligencedelivery of health careeducationgenerative artificial intelligencelarge language models

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Related Experiment Videos

Last Updated: Jun 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

  • Performance was assessed using key metrics: sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and overall accuracy.
  • Confusion matrices were constructed to analyze and compare the AI models' responses.
  • Main Results:

    • Initial assessment showed perfect sensitivity for all chatbots. ChatGPT and Perplexity had the highest initial accuracy (0.80).
    • In a second evaluation round, Gemini 2.5 achieved the highest specificity (0.80) and overall accuracy (0.93).
    • All chatbots demonstrated perfect negative predictive values (1.00) in both rounds, indicating reliable exclusion of incorrect information.

    Conclusions:

    • While initial performance varied, most LLM-based chatbots improved in subsequent evaluations, particularly in specificity and accuracy.
    • Gemini 2.5 emerged as the top-performing chatbot in the second round, demonstrating superior accuracy and specificity.
    • Further research is recommended to explore the full potential and limitations of LLMs in healthcare delivery.