Related Experiment Video
Updated: Jan 17, 2026

Temporomandibular Joint Pain Measurement by Bite Force and Von Frey Filament Assays in Mice
Published on: September 13, 2024
Can Chatbots Provide Accurate and Readable Information for Patients With Temporomandibular Disorders?
Luís Eduardo Charles Pagotto1, Dennys Ramon de Melo Fernandes Almeida2, Thiago de Santana Santos3
1Oral and Maxillofacial Surgeon, Clinical Staff of the Sirio-Libanês Hospital. Hospital Sírio-Libanes, São Paulo, Brazil.
Large language model chatbots offer accurate and empathetic information for temporomandibular disorders (TMDs). Mixture-of-experts models excel in readability and reliable citations, unlike transformer-based models with frequent reference errors.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Natural Language Processing
Background:
- Temporomandibular disorders (TMDs) are prevalent conditions impacting jaw function and patient quality of life.
- Patients with TMDs often struggle to find trustworthy health information.
- Large language models (LLMs) offer potential as educational tools, but concerns about accuracy, readability, empathy, and citation integrity persist.
Purpose of the Study:
- To evaluate the clinical accuracy, empathy, and readability of LLM-based chatbots for patient-friendly TMD queries.
- To assess the authenticity and retrievability of citations provided by LLM-based chatbots.
Main Methods:
- A cross-sectional in silico study utilized 23 standardized TMD questions as prompts for three distinct LLM architectures: GPT-4, Gemini Pro, and DeepSeek-V3.
- Accuracy was determined by oral medicine specialists, empathy by expert scoring, and readability using Flesch-Kincaid metrics.
- Citation reliability was verified against PubMed and other authoritative databases.
Main Results:
- No significant differences in accuracy or empathy were found across chatbot platforms (P = .2).
- The mixture-of-experts transformer (DeepSeek-V3) demonstrated superior readability (Flesch-Kincaid Grade Level = 12.19; P < .001).
- The transformer-based autoregressive model (GPT-4) exhibited the highest rate of hallucinated references (47.2%), compared to multimodal (18.8%) and mixture-of-experts (10.1%) models (P < .001).
Conclusions:
- LLM-based chatbots generally provide accurate and empathetic responses for TMD information.
- The mixture-of-experts transformer excels in readability and citation reliability.
- The high incidence of fabricated references in transformer-based models necessitates human oversight for clinical applications.
Related Concept Videos
Myasthenia Gravis: Overview and Treatment
These antibodies interfere with the function of the nicotinic receptors in three ways: by binding to the receptor and disrupting acetylcholine binding; by causing cross-linking of receptors which...
Myasthenia Gravis: Diagnostic Tests
The edrophonium test is a diagnostic tool for myasthenia gravis. It involves...

