YOLO11-guided swin transformer for molar occlusion classification
Mohamed Hosny1, Bayan Abusafia2,3, Ibrahim A Elgendy4
1Center for Finance and Digital Economy, King Fahd University of Petroleum & Minerals, Dhahran, 31261, Saudi Arabia.
Abstract:
Molar occlusion identification is a fundamental component of orthodontic diagnosis and treatment planning. Conventional assessment using intraoral photographs relies on manual visual inspection, making it time-consuming and susceptible to inter-observer variability. Moreover, existing deep learning (DL)-based approaches either analyze the entire image without explicit anatomical localization or depend on manually defined regions of interest, thereby limiting clinical applicability. Accordingly, this study proposes a DL framework for molar occlusion classification. The proposed framework blends YOLO11-based anatomical localization with a Swin Transformer (Swin-T) for occlusion recognition. YOLO11 performs automatic instance segmentation to identify the molar region and eliminate non-diagnostic background content. The resulting prediction-guided molar crops are then processed by Swin-T, which learns discriminative representations of posterior occlusal relationships through hierarchical shifted-window attention and multiscale feature modeling. A dataset of 1101 intraoral images spanning five occlusion classes was used. The localization stage achieved a Dice coefficient of 96.82%. The developed framework attained classification accuracy of 98.17%, outperforming expert orthodontist assessment. Furthermore, the proposed framework outperformed well-established models, including ResNet50V2, DenseNet201, MobileNetV3-Large, EfficientNetV2-S, and vision transformer (ViT-B/16). Experimental results demonstrated perfect recall for Half Class II, Class III, and Half Class III, while maintaining strong predictive performance for Class I and Class II. Localization and explainability analyses confirmed that the framework consistently attended to clinically relevant posterior occlusal structures, affirming the anatomical plausibility of its predictions. These findings provide a promising foundation for automated molar occlusion assessment and suggest potential for future integration of DL-driven decision support into digital orthodontic workflows.
Related Concept Videos
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
Three-Winding Transformers
In the per-unit equivalent circuit of a grounded Y-Y three-phase...
The Ideal Transformer
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's tangential component...
Motor Unit Stimulation
The latent period of contraction marks the onset of excitation-contraction coupling, when the action potential propagates across the sarcolemma, preparing the muscle fibers for contraction. As the fibers enter the contraction phase, the...
Instrument Transformers
Torque On A Current Loop In A Magnetic Field
Consider a rectangular current-carrying loop containing N turns of wire, placed in a uniform magnetic field. The net force on a current-carrying loop...

