Related Experiment Video
Updated: Jan 10, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Limitations of knowledge competency and error patterns in large language models based on orthodontic licensing
Zhang Ruoyan1,2, Liu Liu1,2, Zhao Qian1,2
1Medical School of Chinese PLA, Beijing, 100853, China.
Deepseek-R1 (DS) significantly outperformed ChatGPT-4 (GPT) on orthodontic licensing exams, achieving higher accuracy. However, both large language models (LLMs) exhibit limitations in clinical reasoning and require clinician verification.
Area of Science:
- Artificial Intelligence in Dentistry
- Medical Education Technology
- Orthodontic Assessment
Background:
- Large language models (LLMs) are increasingly explored for medical applications.
- Evaluating LLM performance in specialized fields like orthodontics is crucial for understanding their utility and limitations.
- Orthodontic licensing examinations provide a standardized benchmark for assessing medical knowledge and clinical reasoning.
Purpose of the Study:
- To assess the performance limitations of two leading LLMs, Deepseek-R1 (DS) and ChatGPT-4 (GPT), in the field of orthodontics.
- To compare LLM performance across different orthodontic knowledge domains and identify common error types.
Main Methods:
- A total of 396 text-based questions from the Chinese National Orthodontic Specialist Licensing Examination were used.
- Questions were categorized by knowledge domains (foundational biomechanics, cross-disciplinary integration, specialized theory, clinical decision-making) and error types (factual, logical, semantic).
- Performance metrics including accuracy and error rates were compared between DS and GPT.
Main Results:
- DS achieved significantly higher overall accuracy (80.3%) compared to GPT (52.3%) (p < 0.001).
- DS also demonstrated superior performance in foundational knowledge (79.8% vs 43.4%) and cross-disciplinary domains (81.0% vs 53.0%).
- Factual errors were most common for both models, while DS showed a higher rate of logical errors.
Conclusions:
- DS exhibits greater proficiency than GPT in orthodontic knowledge assessment via standardized exams.
- Both LLMs demonstrate limitations in specialized orthodontic domains requiring complex clinical reasoning.
- While DS shows promise for AI-assisted orthodontic training and evaluation, clinician oversight remains essential due to persistent errors and domain-specific weaknesses.
Related Concept Videos
Language and Cognition
Teeth
In the bud stage, the tooth germ (an aggregation of cells) starts to form in the developing jawbone. During the cap stage, the tooth germ differentiates into enamel organ, dental papilla, and dental sac, which will later develop into the tooth's enamel, dentin...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Barriers to Effective Communication II
Cultural barriers:
Differences in values, beliefs, religion, knowledge, and tradition can significantly impact communication. Awareness of nonverbal cues is critical, especially when conversing with a patient from a different culture. What appears appropriate in one culture may be inappropriate in another.
Semantic barriers:
As a result of their tendency to use...

