Related Experiment Video
Updated: Jan 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of AI Chatbots on Head and Neck Pathology Board-Style Exam Questions and Guidelines for Responsible Use
Zaid H Khoury1,2, Ahmed S Sultan3,4,5,6
1Department of Oral Diagnostic Sciences & Research, Meharry Medical College School of Dentistry, 1801 Meharry Blvd, Nashville, TN, 37208, USA. ZKhoury@mmc.edu.
Artificial intelligence (AI) chatbots show high accuracy in answering head and neck pathology questions but struggle with citation accuracy and image-based questions. Further research is needed for responsible AI integration in pathology education.
Area of Science:
- Medical Education
- Artificial Intelligence in Pathology
- Head and Neck Pathology
Background:
- The integration of artificial intelligence (AI), including large language models (LLMs) or AI chatbots, into medical education requires thorough evaluation.
- While AI performance has been studied in various medical fields, a gap exists in assessing its utility in head and neck pathology.
- This study addresses the need to evaluate AI chatbots' capabilities in head and neck pathology, a complex subspecialty.
Purpose of the Study:
- To conduct a pilot study evaluating the performance of six AI chatbots on head and neck pathology board-style multiple-choice questions (MCQs).
- To assess both response accuracy and citation accuracy of AI chatbot responses.
- To explore the potential of AI-generated MCQs for pathology education.
Main Methods:
- Twenty board-style MCQs relevant to head and neck pathology were sourced from the public domain.
- Six AI chatbots answered the 20 MCQs, generating a total of 120 responses.
- Responses were evaluated for accuracy of the answer and the accuracy of any provided citations.
Main Results:
- AI chatbots demonstrated high accuracy (85-100%) in answering head and neck pathology MCQs.
- Citation accuracy was found to be poor across the evaluated chatbots.
- Performance on image-based questions was also suboptimal, indicating limitations in visual interpretation.
Conclusions:
- AI chatbots show promise for answering factual head and neck pathology questions but have significant limitations in citation and image-based assessments.
- AI-generated MCQs were found to be at a fundamental level, suitable for introductory pathology learners.
- This pilot study provides initial insights into the limitations of AI in pathology and informs guidelines for responsible AI use in education and practice.
Related Concept Videos
Cardiopulmonary Resuscitation II: ACLS Airway Management
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
