Related Experiment Video
Updated: Mar 3, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Assessing Large Language Models for Early Article Identification in Otolaryngology-Head and Neck Surgery Systematic
Ajibola B Bakare1, Young Lee2, Jhuree Hong3
1Tulane University School of Medicine New Orleans Louisiana USA.
Health Care Science
|March 2, 2026
Summary
Large language models like ChatGPT and Bard show potential for identifying recent articles in Otolaryngology systematic reviews, but inaccuracies necessitate human oversight for now.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Otolaryngology Research
Background:
- Systematic literature reviews are crucial in Otolaryngology-Head and Neck Surgery.
- Initial article identification is a critical and time-consuming step.
- The emergence of advanced AI tools necessitates evaluating their role in research methodologies.
Purpose of the Study:
- To assess the effectiveness of ChatGPT and Bard in the initial article identification phase of systematic reviews.
- To compare the performance of ChatGPT and Bard against established systematic review protocols.
- To determine the accuracy and recall of AI-generated article lists.
Main Methods:
- Replication of three PRISMA-based systematic reviews using ChatGPTv3.5 and Bard.
- Comparison of AI-generated outputs (author, title, year, journal) with original references.
- Cross-referencing AI outputs with medical databases to verify authenticity and measure recall.
Main Results:
- Bard demonstrated higher recall and a broader date range in some reviews.
- ChatGPT-2 showed higher recall and more authentic outputs in another review.
- Both AI models produced inaccuracies but identified relevant articles missed by manual searches.
Conclusions:
- Large language models (LLMs) currently fail to fully replicate systematic review methodologies.
- LLMs show promise in identifying relevant, particularly recent, literature.
- Human-led systematic reviews remain the gold standard, but AI refinement holds future potential.

