Related Experiment Video
Updated: May 31, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Multidimensional evaluation of large language models on the AAP in-service examination: Assessing accuracy,
Prita Abhay Dhaimade1, Robin Henderson2
1Division of Periodontology, Department of Surgical Sciences, School of Dental Medicine, East Carolina University, Greenville, North Carolina, United States of America.
None:
Large language models (LLMs) have demonstrated rapid advancements in natural language understanding and generation, prompting their integration into biomedical research, clinical practice, and professional education. However, systematic evaluation of LLMs in specialty-specific domains such as dentistry and periodontology remains limited, particularly regarding multidimensional performance metrics. This study conducted a comprehensive assessment of commercially available LLMs - GPT-4.0, GPT-5.0, and Claude Sonnet 4.0 - on the American Academy of Periodontology In-Service Examination, focusing on response accuracy, self-assessed confidence calibration, citation validity, and hallucination prevalence. Models were evaluated on the 2024 AAP In-Service Examination (331 questions) using two formats: Full Test (all questions at once) and Individual Question (one at a time). Prompts were standardized; models selected answers, and GPT-5.0 and Claude Sonnet 4.0 also provided confidence ratings and citations. Citation validity was assessed using a human-in-the-loop protocol with expert review. Statistical analyses included chi-square, McNemar's, and logistic regression to assess accuracy, question fatigue, confidence calibration, and citation reliability. LLMs achieved high overall accuracy (78-87%), with the Individual Question format consistently yielding higher scores than Full Test, though differences were not statistically significant. Accuracy was highest in fact-dense domains (biochemistry, physiology, microbiology) and lowest in integrative domains (diagnosis, therapy). Significant question fatigue was observed in GPT-5.0 Full Test mode (OR = 0.997, p = 0.035) but not in Individual Question mode. Confidence scores predicted accuracy, with the strongest calibration in Individual Question mode. Citation analysis revealed frequent hallucinations, mostly critically erroneous, and citation validity was independent of answer accuracy. LLMs show promise as adjunctive tools for periodontal education, but their outputs, especially for complex reasoning and citations require rigorous human review to ensure accuracy and safety.