Related Experiment Video
Updated: Jul 8, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
A compact texture frequency spatial gating network for oral cancer classification from clinical oral images
Mahmudul Haque Rijvi1, Preety Arifa Momotaj2, Syed Mohammed Muhive Uddin3
1Department of Information Technology, Washington University of Science and Technology, Alexandria, VA, USA.
None:
Image-level classification of oral-cavity photographs remains challenging because oral-cancer appearances are heterogeneous and many public datasets lack patient identifiers, acquisition metadata, biopsy confirmation, and lesion-level annotations. This study proposes OCAT-Net-S, a compact hierarchical network for binary classification of Normal versus Oral Cancer from RGB clinical oral images. The model combines a convolutional stem, MBConv blocks, MBConv-SE blocks, and Texture Frequency Spatial Gating (TFSG) blocks to integrate local texture extraction, channel recalibration, and frequency-guided spatial attention within a 3.55 M-parameter architecture. Internal evaluation used the Atef-Kaggle Oral Cancer Images for Classification dataset, containing 1,238 images: 553 Normal and 685 Oral Cancer. Because patient identifiers were unavailable, five-fold stratified group cross-validation with filename-sequential grouping was used as a leakage-reduction proxy and repeated across five seeds. OCAT-Net-S achieved an internal AUC-ROC of 0.956, oral-cancer F1-score of 0.942, and sensitivity of 0.942, compared with AUC-ROC 0.941 for the strongest evaluated baseline, TinyViT-5 M. Validation-only temperature scaling produced a mean Brier score of 0.0419, ECE-15 of 0.0174, and NLL of 0.1798. To reduce overinterpretation from repeated seed runs, fold-level inference was reported, giving an internal AUC-ROC estimate of 0.956 with an approximate 95% CI of 0.940-0.972. External transfer testing was performed without external threshold tuning on two independent settings: SMART-OM Normal-versus-OSCC and the Oral Images Dataset benign-versus-malignant oral-lesion task. AUC-ROC decreased to 0.830 and 0.840, respectively, indicating measurable cross-dataset shift despite partial transferability. Qualitative Grad-CAM visualizations suggested lesion-focused activation patterns, but lesion-level masks were unavailable for quantitative localization validation. Overall, OCAT-Net-S demonstrates compact internal image-level discrimination and limited external transferability; however, the findings do not establish patient-level generalization, prospective clinical validity, clinician-equivalent performance, or deployment readiness.