Related Experiment Video
Updated: May 13, 2026

Mixed Reality Assisted Radical Endoscopic Thyroidectomy
Published on: January 31, 2025
Clinical Applications of Multimodal Artificial Intelligence in Otolaryngology: A State-of-the-Art Review
Ying Jie Li1, Flora Su2, Norbert Banyi3
1Faculty of Medicine, University of British Columbia, Vancouver, British Columbia, Canada.
Objective:
Artificial intelligence (AI) has advanced to simultaneously process visual, auditory, and textual inputs, providing users with "multimodal" AI. Given the clinical integration potential of these tools, otolaryngologists must stay informed. This study reviews current literature on applications of multimodal AI in otolaryngology.
Data Sources:
The MEDLINE, EMBASE, SCOPUS, Cochrane Library, Web of Science, and CINAHL databases.
Review Methods:
Databases were searched from the date of inception to March 4, 2025, following Preferred Reporting Items for Systematic Reviews and Meta-analyses extension for scoping reviews (PRISMA-ScR) guidelines. Studies on any application of multimodal AI in otolaryngology were included.
Conclusions:
Forty-four studies were included, with 55% (24/44) published in 2024 and 18% (8/44) in 2025. Image and text were the most commonly combined modalities (80%, 35/44), with emerging combinations including video with vector data (2%,1/44) and omics with text and/or image (14%, 6/44). Head and neck cancer was the most common subspecialty of focus (75%, 33/44), followed by general ear, nose, and throat (ENT) (11%, 5/44). All studies applied the models for clinical education (9%, 4/44) or decision support (91%, 40/44), assessing performance in areas such as board-style examination performance (accuracy: 37%-86%) or disease classification and prognostication (area under the receiver operating characteristic curve [AUC] 0.65-0.96). However, most studies were limited to small, single-institution samples and lacked prospective validation. Model error, data set bias, and language limitations underscore the need for further refinement.
Implications For Practice:
The application of multimodal large language models (LLMs) in otolaryngology is rapidly expanding. Clinicians must understand both the capabilities and limitations of these systems. Rigorous validation and ethical oversight will be essential to ensure the safe, equitable, and effective adoption in otolaryngologic care.