Related Experiment Video
Updated: Jun 7, 2025

05:49
Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
771
Preparing for downstream tasks in artificial intelligence for dental radiology: a baseline performance comparison of
Fara A Fernandes1,2, Mouzhi Ge2, Georgi Chaltikyan2
1Department of Information and Communication Technology, University of Agder (UiA), 4879 Grimstad, Norway.
Dento Maxillo Facial Radiology
|November 20, 2024
Summary
The Vision Transformer (ViT) and convolutional neural network (CNN) show comparable performance in classifying dental radiographic images. While the gated multilayer perceptron (gMLP) performed slightly lower, different AI architectures offer unique advantages for specific dental imaging tasks.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Deep Learning for Dental Diagnostics
- Radiographic Image Analysis
Background:
- Convolutional Neural Networks (CNNs) are widely used for medical image classification.
- Emerging deep learning models like Vision Transformers (ViTs) and gated Multilayer Perceptrons (gMLPs) offer alternative architectures.
- Evaluating these models for dental radiographic analysis is crucial for advancing diagnostic accuracy.
Purpose of the Study:
- To compare the classification performance of CNN, ViT, and gMLP models.
- To assess the effectiveness of these deep learning architectures on diverse dental radiographic tasks.
- To determine the potential of ViT and gMLP as alternatives to CNNs in dental imaging.
Main Methods:
- Utilized retrospective 2D radiographic images from cone beam computed tomographic volumes.
- Trained and evaluated CNN, ViT, and gMLP classifiers on four distinct dental cases.
- Calculated performance metrics including sensitivity, specificity, accuracy, F1-score, and AUC-ROC.
Main Results:
- ViT and CNN demonstrated comparable performance across all classification tasks, with accuracies ranging from 0.71 to 0.99.
- gMLP showed slightly lower performance (accuracy 0.65-0.98) compared to ViT and CNN.
- Area Under the Curve (AUC) values were high for all models, ranging from 0.73 to 1.00.
Conclusions:
- ViT and gMLP models perform comparably to the current state-of-the-art CNNs in dental radiographic classification.
- Performance variations among architectures highlight task-specific strengths.
- The capabilities of different deep learning architectures can be leveraged for optimized dental image analysis.

