Related Experiment Video
Updated: Aug 15, 2026

09:10
Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Vision-Language Model as a 'Zero-Shot' Assistant for Evaluating Condylar Osseous Changes in Cone-beam Computed
Ke Chen1, Andrew Zhang2, Xianju Xie1
1Department of Orthodontics, Beijing Stomatological Hospital, Capital Medical University, Beijing, China.
International Dental Journal
|August 13, 2026
Summary
Vision-Language Models (VLMs) show promise for detecting condylar changes on CBCT scans. Gemini-3 demonstrated high accuracy, offering potential as a preliminary AI screening tool for general practitioners.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Oral and Maxillofacial Radiology
- Computer-Aided Diagnosis
Background:
- Interpreting cone-beam computed tomography (CBCT) for condylar osseous changes presents diagnostic challenges for general dental practitioners.
- The need for advanced, accessible tools to aid in the detection of temporomandibular joint abnormalities is growing.
Purpose of the Study:
- To evaluate the 'zero-shot' diagnostic performance of Vision-Language Models (VLMs) in identifying condylar abnormalities on CBCT images.
- To assess the utility of VLMs as AI assistants for general practitioners in detecting osseous changes in the mandibular condyle.
- To analyze the quality of AI-generated diagnostic reports using a standardized framework.
Main Methods:
- Analysis of 72 CBCT images (EHPN study) for internal validation and 70 images (MMDental dataset) for external validation, with dual-radiologist consensus as ground truth.
- Three frontier VLMs (Gemini-3, GPT-5.2, Qwen3-VL) were tested on representative sagittal slices without prior fine-tuning.
- Performance metrics included diagnostic accuracy, sensitivity, specificity, and balanced accuracy; report quality was assessed using the QAMAI framework, adhering to STARD-AI, CLAIM, and TRIPOD-LLM guidelines.
Main Results:
- Gemini-3 achieved the highest diagnostic accuracy (90.3% internal, 90.0% external validation), significantly outperforming GPT-5.2 (75.0%) and Qwen3-VL (55.6%).
- Gemini-3 also excelled in generating structured diagnostic reports, scoring higher across QAMAI dimensions for accuracy, clarity, and clinical utility.
- All evaluated VLMs showed a conservative diagnostic tendency.
Conclusions:
- Vision-Language Models, particularly Gemini-3, demonstrate effective 'zero-shot' capabilities for detecting condylar changes and generating high-quality diagnostic reports.
- These AI models show potential as preliminary, training-free screening aids in oral and maxillofacial radiology, enhancing diagnostic workflows for general practitioners.
- Further multicenter validation is recommended to confirm clinical utility, with cloud-based VLMs offering accessible screening assistance via simple web interfaces.
Keywords:
Artificial IntelligenceCone-beam computed tomographyLLMTemporomandibular JointVision-language model
