Comparing Large Language Models' Performances on Otolaryngology Knowledge Assessment Questions

Ryan Cook1, Abner Kahan1, Thomas Scharfenberger1

  • 1Albert Einstein College of Medicine, Bronx, New York, United States.

Summary

This study tested large language models (LLMs) on otolaryngology knowledge. Top models achieved around 76% accuracy, indicating a plateau for general LLMs in specialized medical fields.