Related Experiment Video
Updated: Sep 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the accuracy and usability of artificial intelligence-based language models in responding to common
Sajad Jahantigh1, Reza Amid1, Anahita Moscowchi2
1School of Dentistry, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Background:
Artificial intelligence-powered large language models (LLMs) are increasingly used by patients seeking quick information regarding dental and medical problems. Despite their growing popularity, concerns remain regarding the accuracy, clarity, and clinical usefulness of LLM-generated responses. This study aimed to comparatively evaluate the accuracy and usability of widely used LLMs in responding to frequently asked periodontal questions.
Methods:
In this analytical-comparative study, 15 commonly asked periodontal questions were selected from real patient encounters and administered to ten LLMs in both English and Persian. Responses were evaluated independently and blindly by two board-certified periodontists; disagreements were resolved by discussion or third-reviewer. Each response was rated on a 5-point Likert scale across six criteria: correlation with the question, adequacy, comprehensiveness, clarity/readability, usability for individuals with limited scientific literacy, and scientific accuracy. Statistical analyses included descriptive statistics, Mann-Whitney U test, Kruskal-Wallis test with Bonferroni correction, and univariate ANOVA. Significance level was set at p < 0.05.
Results:
Significant differences were observed among LLMS (χ2 = 294.78, p < 0.001). ChatGPT-4.5 with deep search enabled achieved the highest scores across most criteria, while ChatGPT-4o generally underperformed compared with other LLMs. English responses scored significantly higher than Persian (mean 4.25 vs. 3.96; U = 333,058, p < 0.001). Among evaluated criteria, correlation had the highest mean score (4.81), whereas usability for individuals with limited scientific literacy had the lowest score (3.80). LLMs performed significantly better on treatment/prevention questions than on diagnosis/pathogenesis questions. Only ChatGPT-4.5 (deep search) and Grok-3 (deep search) provided citations, with ChatGPT-4.5 referencing more specialized dental sources.
Conclusions:
LLMs show variable performance in answering frequent periodontal patients' questions, influenced by model architecture, language, question type, and evaluative domain. Although advanced configurations such as deep-search ChatGPT-4.5 offer highly accurate, comprehensive, and well-referenced responses, significant limitations remain particularly in Persian language output and in usability for patients with limited scientific literacy. LLMs may serve as complementary tools for patient communication, but should not replace professional advice. Further development and language-specific optimization are needed to improve global accessibility and reliability.
Key Points:
While AI tools provide highly readable and generally accurate periodontal information, they often lack the clinical nuance required for personalized risk assessment and complex treatment planning. Practitioners should proactively discuss AI-generated health information with patients to address potential oversimplifications and ensure online advice is integrated into a professional, evidence-based care plan. As patients increasingly rely on AI for initial health guidance, practitioners must act as the authoritative "filter" to verify that online information remains clinically safe, accurate, and aligned with the patient's specific treatment goals.
Plain Language Summary:
Many people now turn to artificial intelligence for help with their dental questions. This study found that these tools vary in how well they answer questions about gum disease. While the most advanced models (such as ChatGPT-4.5) provide accurate, high-quality information, there are two main limitations: first, performance is significantly better in English than in Persian; second, many explanations are too technical and hard for the average person to understand. Overall, while AI can be a helpful guide, it should never replace a professional consultation with a dentist.