Related Experiment Video
Updated: Jan 10, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Limitations of knowledge competency and error patterns in large language models based on orthodontic licensing
Zhang Ruoyan1,2, Liu Liu1,2, Zhao Qian1,2
1Medical School of Chinese PLA, Beijing, 100853, China.
Objectives:
This study aimed to evaluate the limitations of large language models (LLMs) in orthodontics by comparing their performance on licensing exam questions categorized by knowledge domains and error types.
Materials And Methods:
Deepseek-R1(DS) and ChatGPT-4(GPT) were evaluated using 396 text-based questions from the Chinese National Orthodontic Specialist Licensing Examination. Questions were classified through dual taxonomies: (1) "knowledge domains" including foundational biomechanical principles, cross-disciplinary medical integration, specialized orthodontic theory, and clinical decision-making skills; (2) "error types" including factual inaccuracies, logical deficits, and semantic misinterpretations.
Results:
DS demonstrated significantly higher overall accuracy than GPT (80.3% vs 52.3%, p < 0.001), with statistically significant differences in foundational knowledge (79.8% vs 43.4%) and cross-disciplinary domains (81.0% vs 53.0%). Factual errors were predominant in both models (DS:57.7%, GPT:69.3%), though DS exhibited higher logical error rates (24.4% vs 16.4%).
Conclusions:
While DS outperforms GPT in general orthodontic knowledge assessment, both models show limitations in specialized domains requiring clinical reasoning.
Clinical Relevance:
The superior performance of DS in standardized exams suggests potential for AI-assisted decision support in orthodontic training and licensing evaluation. However, the persistent factual errors and domain-specific limitations highlight the necessity for clinician verification in real-world applications. Integrating domain-specific knowledge refinement with logical reasoning modules could enhance LLMs' clinical utility in orthodontic practice.
Related Concept Videos
Language and Cognition
Teeth
In the bud stage, the tooth germ (an aggregation of cells) starts to form in the developing jawbone. During the cap stage, the tooth germ differentiates into enamel organ, dental papilla, and dental sac, which will later develop into the tooth's enamel, dentin...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Barriers to Effective Communication II
Cultural barriers:
Differences in values, beliefs, religion, knowledge, and tradition can significantly impact communication. Awareness of nonverbal cues is critical, especially when conversing with a patient from a different culture. What appears appropriate in one culture may be inappropriate in another.
Semantic barriers:
As a result of their tendency to use...

