Related Experiment Video
Updated: Jan 17, 2026

02:37
Robot-Assisted Transcanal Endoscopic Ear Surgery for Congenital Cholesteatoma
Published on: December 15, 2023
1.3K
Generative Artificial Intelligence Methodology Reporting in Otolaryngology: A Scoping Review
Isaac L Alter1, Karly Chan1, Katerina Andreadis2
1Department of Otolaryngology-Head and Neck Surgery, Sean Parker Institute for the Voice, Weill Cornell Medical College, New York, New York, USA.
The Laryngoscope
|September 25, 2025
Summary
Reporting of large language model (LLM) methods in otolaryngology is inconsistent, hindering reproducibility. Studies often omit crucial details on prompt engineering and model parameters, impacting the generalizability of LLM applications in OHNS research.
Area of Science:
- Otolaryngology-Head and Neck Surgery (OHNS)
- Artificial Intelligence
- Medical Informatics
Background:
- Large language models (LLMs) show promise in OHNS research.
- Inconsistent reporting of LLM methodology, particularly prompt engineering and model parameters, poses a challenge to study reproducibility.
- There is a need to assess the quality of methodological reporting in LLM-focused OHNS literature.
Purpose of the Study:
- To critically review the methodological reporting and quality of studies utilizing LLMs within the field of OHNS.
- To identify gaps in the reporting of essential LLM implementation details.
Main Methods:
- A systematic search of multiple databases (PubMed, Embase, Web of Science, etc.) was conducted in October 2024.
- Two independent reviewers performed abstract and full-text review, including data extraction for all primary studies using LLMs in OHNS.
- 117 studies were included from 925 retrieved abstracts.
Main Results:
- All included studies utilized ChatGPT; only 16.2% incorporated additional LLMs.
- Prompt reporting was incomplete: 46.2% provided prompt quotations, 76.9% reported prompt numbers but few rationalized them (6.8%), and 23.9% reported runs per prompt.
- While 73.5% described prompt development, only 11.1% explained design decisions, and 6.0% reported prompt testing. Reporting quality did not improve over time.
Conclusions:
- LLM literature in OHNS is exploring valuable areas but suffers from variable methodological reporting completeness.
- Incomplete reporting significantly limits the generalizability of findings from these studies.
- Dissemination and enforcement of best practices for LLM reporting are recommended for researchers and journals.

