Related Experiment Video
Updated: Sep 16, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Information Quality and Semantic Stability of AI Chatbot Responses to Complete Denture Questions: A Turkish-English
Sinan Coşkun1, Fatma Nur Karaman1, Hüseyin Ardıl Uytun1
1Department of Prosthodontics, Faculty of Dentistry, Ege University, Bornova 35040, İzmir, Türkiye.
Abstract:
Background/Objectives: Artificial intelligence (AI)-based chatbots are increasingly used as sources of health information, yet the quality and semantic stability of their responses may vary across systems, languages, repeated queries, and prompt formulations. This study compared expert-rated general information quality and embedding-based semantic stability of responses generated by seven AI chatbot systems to expert-derived, patient-oriented complete denture questions in Turkish and English. Methods: Twenty-two complete denture-related questions were submitted to seven chatbot systems on three study days across scheduled morning, afternoon, and evening sessions. Two additional semantically equivalent Turkish variants were generated for each original question. Five prosthodontists assessed the original Turkish and English responses using the 5-point Global Quality Score (GQS). Semantic stability was quantified using the multilingual SentenceTransformer checkpoint paraphrase-multilingual-MiniLM-L12-v2 and cosine similarity. To incorporate complete responses, responses were divided into non-overlapping token-based chunks, chunk embeddings were combined into one normalized full-response representation, and the 36 pairwise similarities from the nine repeated responses were averaged to one stability estimate per question-model-language/variant condition. Linear mixed-effects models with question-level clustering were used, with Bonferroni adjustment and partial eta squared effect sizes with 95% confidence intervals. Results: AI model significantly affected embedding-based semantic stability (p < 0.001; ηp2 = 0.820, 95% CI 0.795-0.837) and GQS (p < 0.001; ηp2 = 0.249, 95% CI 0.152-0.316). Grok had the numerically highest overall semantic stability mean (0.952 ± 0.015), whereas GPT-4o had the numerically highest overall GQS (4.48 ± 0.85). English responses showed higher overall GQS than the original Turkish responses (3.96 ± 0.94 vs. 3.66 ± 1.23; p = 0.004) and higher embedding-based similarity than the Turkish conditions overall (p < 0.001). No significant overall effect of scheduled study day was observed (p = 0.504; ηp2 = 0.001), whereas semantic stability differed across scheduled query sessions (p < 0.001; ηp2 = 0.055), with means of 0.863 ± 0.106, 0.874 ± 0.092, and 0.841 ± 0.140 for morning, afternoon, and evening sessions, respectively. Inter-rater reliability was moderate for English GQS ratings (ICC = 0.675, 95% CI 0.581-0.752) and good for Turkish ratings (ICC = 0.771, 95% CI 0.687-0.833). Conclusions: The evaluated chatbot systems differed in expert-rated information quality and full-response embedding-based semantic stability. English responses showed higher overall values than Turkish responses, although cross-language measurement effects of the embedding model cannot be excluded. Differences across scheduled query sessions should be interpreted as run-to-run output variability rather than intrinsic temporal behavior. These findings characterize comparative chatbot performance under the tested conditions but do not establish clinical accuracy, safety, or patient education effectiveness.