Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Assessment of Airway, Skin Color, and Use of Accessory Muscles01:30

Assessment of Airway, Skin Color, and Use of Accessory Muscles

A thorough assessment of respiratory health is paramount in clinical settings to identify and manage respiratory distress and ensure adequate oxygenation. This article elaborates on the critical aspects of respiratory evaluation, including airway assessment, skin color examination, and the observation of accessory muscle use, which are integral to effectively diagnosing and managing patients with respiratory conditions.
Introduction
The initial evaluation of a patient's respiratory system...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Porcine Model of Intervertebral Disc Injury Recapitulates Human Discogenic Pain Via Notochordal Cell Loss and Pain-Inducing Nucleus Pulposus Cell Emergence.

JOR spine·2026
Same author

Biosynthesis of antibacterial zinc oxide nanoparticles from endophytic Streptomyces werraensis.

Scientific reports·2026
Same author

Clinical and electrophysiological effects of cathodal high-definition transcranial direct current stimulation to the right superior temporal sulcus in psychosis spectrum disorders: a pilot proof of concept study.

International review of psychiatry (Abingdon, England)·2026
Same author

Single-cell transcriptomics-informed induced pluripotent stem cells differentiation to tenogenic lineage.

eLife·2026
Same author

Quantitative Assessment of Kidney Interstitial Fibrosis Using an Intrarenal Optical Spectroscopy.

Kidney360·2026
Same author

Telehealth for Sexual and Reproductive Healthcare: Evidence Map of Effectiveness, Patient and Provider Experiences and Preferences, and Patient Engagement Strategies.

Clinics and practice·2026

Related Experiment Video

Updated: May 31, 2026

Assessment and Communication for People with Disorders of Consciousness
07:37

Assessment and Communication for People with Disorders of Consciousness

Published on: August 1, 2017

Large Language Model Responses to Common Otolaryngological Questions: Evaluating Accuracy, Appropriateness,

Mitali Sakharkar1, Parth Jalihal1, Sarah Chang1

  • 1Boston University Chobanian and Avedisian School of Medicine, Boston, Massachusetts, USA.

Otolaryngology--Head and Neck Surgery : Official Journal of American Academy of Otolaryngology-Head and Neck Surgery
|May 29, 2026
PubMed
Summary

Artificial Intelligence (AI) tools like ChatGPT and Claude often provide inaccurate, jargon-filled answers to otolaryngology questions and frequently fabricate references, indicating a need for caution in clinical use.

Keywords:
Artificial IntelligenceChatGPTLarge Language Modelsmedical misinformationpatient resources

More Related Videos

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)
12:43

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)

Published on: February 21, 2011

Related Experiment Videos

Last Updated: May 31, 2026

Assessment and Communication for People with Disorders of Consciousness
07:37

Assessment and Communication for People with Disorders of Consciousness

Published on: August 1, 2017

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)
12:43

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis (ALS)

Published on: February 21, 2011

Area of Science:

  • Medical Informatics
  • Artificial Intelligence in Healthcare
  • Otolaryngology

Background:

  • Artificial Intelligence (AI) and Large Language Models (LLMs) are increasingly used in medicine.
  • Concerns exist regarding the accuracy and reliability of AI-generated medical content, particularly LLM-generated references.
  • Otolaryngology is an area where AI integration is growing, necessitating evaluation of these tools.

Purpose of the Study:

  • To evaluate the accuracy, appropriateness, readability, and reference hallucination of ChatGPT and Claude for otolaryngologic patient questions.
  • To compare the performance of two leading LLMs in a specific medical domain.

Main Methods:

  • A prospective observational study was conducted at an academic tertiary care center.
  • Thirty-six otolaryngologic questions were posed to ChatGPT 4.0 Plus and Claude twice for reproducibility.
  • Responses were independently assessed for accuracy by two otolaryngologists, with readability measured by Flesch Reading Ease (FRE) and reference validity analyzed for hallucinations.

Main Results:

  • Claude (FRE 25.2) and ChatGPT (FRE 47) demonstrated varying readability scores, with Claude scoring higher for patient comprehension (4.68/5 vs. 3.60/5).
  • Claude achieved higher accuracy (4.42/5) compared to ChatGPT (3.81/5).
  • Both LLMs frequently hallucinated references (over 50%), with irrelevant or poorly formatted citations, and responses often contained jargon and lacked clinical prioritization.

Conclusions:

  • ChatGPT and Claude frequently generate partially inaccurate, jargon-filled responses to otolaryngologic patient queries.
  • Both models struggle with consistently providing valid references, with significant hallucination rates observed.
  • The findings underscore the necessity for improved understanding and regulation of LLM limitations in clinical and patient-facing medical applications.