Related Experiment Video
Updated: Jan 17, 2026

Temporomandibular Joint Pain Measurement by Bite Force and Von Frey Filament Assays in Mice
Published on: September 13, 2024
Can Chatbots Provide Accurate and Readable Information for Patients With Temporomandibular Disorders?
Luís Eduardo Charles Pagotto1, Dennys Ramon de Melo Fernandes Almeida2, Thiago de Santana Santos3
1Oral and Maxillofacial Surgeon, Clinical Staff of the Sirio-Libanês Hospital. Hospital Sírio-Libanes, São Paulo, Brazil.
Background:
Temporomandibular disorders (TMDs) are common musculoskeletal and neuromuscular conditions that impair jaw function and quality of life. Patients often lack access to reliable health information. Large language models (LLMs) have introduced chatbots as potential educational tools, yet concerns remain regarding accuracy, readability, empathy, and citation integrity.
Purpose:
This study evaluated whether LLM-based chatbots can provide clinically accurate, empathic, and readable responses to patient-friendly questions about TMDs and whether their cited references are authentic.
Study Design, Setting, Sample:
This cross-sectional in silico study was conducted in March 2025. Twenty-three standardized TMD-related questions were used as prompts for each chatbot.
Predictor/Exposure/Independent Variable:
The predictor variable was the chatbot platform, reflecting distinct LLM architectures: GPT-4 (transformer-based autoregressive model, OpenAI), Gemini Pro (multimodal transformer, Google), and DeepSeek-V3 (mixture-of-experts transformer, DeepSeek).
Main Outcome Variables:
Accuracy was defined as the proportion of responses judged clinically correct by two board-certified oral medicine specialists. Empathy was assessed by expert scoring of tone. Readability was determined with Flesch-Kincaid Reading Ease and Grade Level. Citation reliability was assessed by verifying whether references were authentic and retrievable in PubMed or other authoritative databases.
Covariates:
No formal covariates were included; exploratory correlations between variables were performed.
Analyses:
Descriptive statistics, 1-way analysis of variance with Tukey's post hoc tests, Pearson correlation, and χ2 tests were performed. Statistical significance was set at P < .05.
Results:
No statistically significant differences were observed in accuracy (P = .2) or empathy (P = .2). The mixture-of-experts transformer provided the most readable content (Flesch-Kincaid Reading Ease = 28.47; Flesch-Kincaid Grade Level = 12.19; P < .001). The transformer-based autoregressive model produced the highest proportion of hallucinated references (47.2%), compared with the multimodal transformer (18.8%) and the mixture-of-experts transformer (10.1%) (P < .001). A weak positive correlation was found between accuracy and readability (r = 0.27; P = .03), with no correlation between accuracy and empathy.
Conclusions And Relevance:
While all LLM-based chatbots delivered generally accurate and empathetic responses, the mixture-of-experts transformer outperformed others in readability and citation reliability. The high rate of hallucinated references in the transformer-based autoregressive model underscores the need for human oversight in clinical applications.
Related Concept Videos
Myasthenia Gravis: Overview and Treatment
These antibodies interfere with the function of the nicotinic receptors in three ways: by binding to the receptor and disrupting acetylcholine binding; by causing cross-linking of receptors which...
Myasthenia Gravis: Diagnostic Tests
The edrophonium test is a diagnostic tool for myasthenia gravis. It involves...

