Related Experiment Video
Updated: Mar 29, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Evaluating Large Language Models for Diagnostic Accuracy and Health Information Quality in Oral Mucosal Diseases
Melisa Iacob1, Ayham Qawas2, Ramesh Balasubramaniam2
1Toothbuds, Perth, WA 6062, Australia.
None:
Background: Multimodal large language model (MLLM)-based systems capable of generating health-related information and diagnostic suggestions are increasingly used for health information retrieval; however, their accuracy, readability, and quality in oral healthcare remain unclear. Oral mucosal diseases comprise a heterogeneous group of conditions affecting the oral lining, ranging from benign and reactive lesions to potentially malignant and malignant disorders. Objective: This study evaluated and compared the diagnostic performance, readability, and information quality of MLLMs with traditional search engines included as comparator platforms, in diagnosing oral mucosal diseases. Methods: A cross-sectional observational study was conducted using 100 validated oral mucosal case scenarios representing benign, malignant, potentially malignant, infectious, and reactive oral lesions. Each scenario was entered into ChatGPT 3.5, ChatGPT 4.5 (Plus), Microsoft Copilot (smart), Grok (xAI), Claude (Sonnet 4.5), DeepSeek v3.1, and search engines Google, Bing, and Yahoo. Diagnostic accuracy, Positive Predictive Value (PPV), and Negative Predictive Value (NPV) were compared against reference diagnoses. Information quality was assessed using the DISCERN tool, and readability was evaluated using Flesch-Kincaid Reading Ease (FRES) and Grade Level (FKGL) scores. Statistical analyses included Cochran's Q and McNemar tests (p < 0.05). Results: ChatGPT 4.5 demonstrated the highest overall diagnostic accuracy (88.5%), PPV (92%), and NPV (88%), followed by DeepSeek v3.1 and Claude (Sonnet 4.5). Traditional search engines performed poorly (accuracy 18-55%). MLLMs achieved higher DISCERN scores (2.84-3.20) but lower readability (FKGL = 11-14) than search engines (FKGL = 6-7). No platform met the recommended sixth-grade reading level for consumer health information. Conclusions: MLLMs, particularly ChatGPT Plus (GPT-4.5), outperformed conventional search engines in diagnostic accuracy and content quality but produced complex, less-readable text. Future AI development should prioritise improving clinical accuracy alongside readability and transparency to ensure equitable access to reliable oral health information.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025