Related Experiment Video
Updated: Aug 31, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
Multimodal transformer-enhanced deep learning for accurate assessment of mandibular third molar-inferior alveolar
Rasool Esmaeilyfard1, Nika Shahbeyk1, Setare Sanchooli2
1Department of Computer Engineering and Information Technology, Shiraz University of Technology, Shiraz, Iran.
Background:
Accurate preoperative assessment of the spatial relationship between mandibular third molars (M3Ms) and the inferior alveolar canal (IAC) is essential to minimize the risk of nerve injury. While panoramic radiographs (PRs) are routinely used as a first-line modality, their two-dimensional nature limits diagnostic reliability, often necessitating confirmatory cone-beam computed tomography (CBCT).
Methods:
We developed a dual-stream, transformer-enhanced multimodal (panoramic radiography + selected CBCT slices) deep learning framework designed as a clinician-in-the-loop system. The model jointly processes paired panoramic radiographs and operator-selected CBCT slices to simulate a realistic specialist consultation. Four widely used high-performance backbones (ConvNeXt, Swin Transformer, ResNet-50, and VGG16) were systematically evaluated on a dataset of 250 paired PR-CBCT cases. Model performance was assessed using F1-score, precision, recall, and area under the ROC curve (AUC). Model interpretability was explored using SHAP-based explainability analysis.
Results:
Among the evaluated architectures, the ConvNeXt-based model achieved the highest overall performance, with a macro F1-score of 0.9478, precision of 0.9515, and recall of 0.9468. The learned embeddings showed good class separability, and explainability analyses highlighted anatomically meaningful regions consistent with established radiographic risk signs.
Conclusion:
The proposed multimodal 2D framework demonstrates competitive and robust classification performance. By integrating expert slice selection with automated feature extraction, the model functions as a standardized decision-support tool, aimed at reducing diagnostic variability in complex boundary cases where visual interpretation alone may be subjective.
