Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Introduction to Language of Pathophysiology l01:25

Introduction to Language of Pathophysiology l

239
Pathophysiology investigates how biological mechanisms—typically starting at the cellular level—disrupt normal bodily functions. It bridges anatomy and physiology to explain the progression of disease. With this foundation, it is important to understand the following key terms used to describe disease processes: Diagnosis:The process of identifying a disease using clinical evaluation, including signs (objective evidence like rashes), symptoms (subjective experiences like...
239
Introduction to Language of Pathophysiology ll01:17

Introduction to Language of Pathophysiology ll

69
This lesson explores key terms that describe how diseases progress, their outcomes, and their distribution in populations.Diagnostic tests identify diseases and monitor treatment. These include blood and urine tests, biopsies, imaging (X-ray, MRI), and detection of infectious agents.Remission is a reduction or disappearance of symptoms.Exacerbation refers to the worsening of symptoms, such as increased wheezing during an asthma attack.A precipitating factor triggers an acute episode, while a...
69

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Long-Term Electromyographic Monitoring of the Stapedius Reflex via Implanted Electrodes in Sheep: Toward Objective Autonomous Cochlear Implant Fitting.

Sensors (Basel, Switzerland)·2026
Same author

Not all acute peripheral facial palsy in children recovers completely: A telediagnostic follow-up study.

European journal of paediatric neurology : EJPN : official journal of the European Paediatric Neurology Society·2026
Same author

Toxin gene truncation via IS1132 in non-toxigenic toxin gene-bearing (NTTB) Corynebacterium diphtheriae.

International journal of medical microbiology : IJMM·2026
Same author

[Bilateral low-frequency hearing loss and tinnitus following spinal anesthesia during a cesarean section].

HNO·2026
Same author

Poly(I:C) Lipoamino Bundle LNPs Induce Tumor Cytotoxicity and Immune Activation with Enhanced Efficacy by Survivin Silencing.

International journal of molecular sciences·2026
Same author

Blood plasma and oral rinse liquid profiling for human papillomavirus in head and neck cancer - Unmasking false-positive p16 tissue cases and tracking disease dynamics.

Journal of translational medicine·2026

Related Experiment Video

Updated: May 5, 2026

Learning Modern Laryngeal Surgery in a Dissection Laboratory
07:30

Learning Modern Laryngeal Surgery in a Dissection Laboratory

Published on: March 18, 2020

7.9K

Harnessing advanced large language models in otolaryngology board examinations: an investigation using python and

Cosima C Hoch1, Paul F Funk2, Orlando Guntinas-Lichius2

  • 1Department of Otolaryngology, Head and Neck Surgery, TUM School of Medicine and Health, Technical University of Munich (TUM), Ismaningerstrasse 22, 81675, Munich, Germany. cosima.chiara.hoch@tum.de.

European Archives of Oto-Rhino-Laryngology : Official Journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : Affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery
|April 25, 2025
PubMed
Summary

Advanced AI language models show promise for otolaryngology exams, but GPT-3.5 Turbo

Keywords:
APIsBoard examinationsLarge language modelsMedical AI integrationOtolaryngology educationPython programming

More Related Videos

Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
01:00

Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing

Published on: December 1, 2023

425
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

462

Related Experiment Videos

Last Updated: May 5, 2026

Learning Modern Laryngeal Surgery in a Dissection Laboratory
07:30

Learning Modern Laryngeal Surgery in a Dissection Laboratory

Published on: March 18, 2020

7.9K
Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
01:00

Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing

Published on: December 1, 2023

425
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

462

Area of Science:

  • Artificial Intelligence in Medicine
  • Medical Education Technology
  • Otolaryngology Diagnostics

Background:

  • Large language models (LLMs) are increasingly evaluated for specialized medical applications.
  • Assessing LLM performance on board certification exams is crucial for understanding their utility.
  • Previous evaluations of GPT-3.5 Turbo provide a baseline for longitudinal performance tracking.

Purpose of the Study:

  • To evaluate the capabilities of advanced LLMs (GPT-4 variants, Gemini, Claude) on otolaryngology board exam questions.
  • To assess the longitudinal performance changes of GPT-3.5 Turbo over one year.
  • To compare the accuracy of different LLM series in a specialized medical domain.

Main Methods:

  • A question bank of 2,576 otolaryngology board certification questions was used.
  • Eleven LLMs, including GPT-3.5 Turbo, GPT-4 variants, Gemini, and Claude, were tested.
  • Questions were administered via APIs using Python scripts for data collection.

Main Results:

  • GPT-4o achieved the highest accuracy, particularly in allergology and head and neck tumors.
  • Claude models were competitive but generally underperformed GPT-4 variants.
  • GPT-3.5 Turbo showed a significant decline in accuracy compared to the previous year.
  • Single-choice questions yielded higher accuracy than multiple-choice questions across all models.

Conclusions:

  • Newer LLMs demonstrate significant potential for specialized medical content and certification.
  • The performance decline in GPT-3.5 Turbo highlights the need for continuous LLM evaluation.
  • Ongoing optimization and efficient API use are essential for advancing LLMs in medical education.