Related Experiment Video
Updated: Aug 13, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
ENTGPT: Applying Large Language Models to Systematic Review Screening With the Novel STARR Protocol
Akash Kapoor1, Ben Baranker1, Isaac Alter1
1Department of Otolaryngology-Head and Neck Surgery, Columbia University Vagelos College of Physicians and Surgeons, New York-Presbyterian/Columbia University Irving Medical Center, New York, New York, USA.
The Laryngoscope
|August 12, 2026
Summary
Large language models (LLMs) show promise for streamlining systematic literature reviews (SLRs) in otolaryngology. ENTGPT, using the STARR protocol, achieved 99.87% accuracy, significantly improving efficiency.
Area of Science:
- Otolaryngology research
- Artificial intelligence in medicine
- Systematic literature reviews
Background:
- Systematic literature reviews (SLRs) are crucial but highly time-intensive.
- Large language models (LLMs) show potential for clinical knowledge but require validation for complex tasks like SLRs.
- Evidence for LLM performance in otolaryngology SLRs is limited.
Purpose of the Study:
- To evaluate an LLM's (ENTGPT, based on GPT-4o) performance in screening articles for otolaryngology SLRs.
- To assess the efficacy of the novel Screening of Title and Abstracts, Reevaluation, and Full-text review (STARR) protocol.
- To compare LLM performance against human reviewers.
Main Methods:
- ENTGPT was compared to two human reviewers using traditional and STARR protocols.
- The LLM screened 850 articles, utilizing inclusion/exclusion criteria, titles, abstracts, and full texts.
- Article inclusion/exclusion decisions were compared between ENTGPT and human reviewers.
Main Results:
- ENTGPT with the STARR protocol achieved 99.87% accuracy, 100% specificity, and 95% sensitivity.
- Using the traditional protocol, ENTGPT's sensitivity dropped to 35% but accuracy remained high at 99.47%.
- The STARR protocol significantly enhanced LLM sensitivity compared to the traditional approach.
Conclusions:
- ENTGPT accurately replicated human reviewers in article selection and data extraction for otolaryngology SLRs.
- The STARR protocol enhances LLM performance in SLR article screening.
- LLMs like ENTGPT can potentially save significant time and resources in conducting SLRs.