Related Experiment Video
Updated: Sep 26, 2026

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
MonuSegFormer: A Hybrid Swin-Transformer Architecture for Semantic Segmentation of Moroccan Cultural Heritage
Ouail Choukhairi1, Mouad Choukhairi1, Ali Choukri1
1Laboratory of Computer Science Research (LaRI), Department of Computer Science, Ibn Tofail University, Kénitra 14000, Morocco.
Abstract:
Semantic segmentation of architectural elements in cultural heritage sites lies at the intersection of computer vision and digital preservation. Moroccan historical monuments spanning mosques, madrasas, royal gates (babs), and mausoleums across Fez, Rabat, Marrakech, Meknes, and Tetouan present unique challenges, including extreme texture ambiguity between weathered wall surfaces and background, pronounced multi-scale variation from individual window and door openings to full rooftop surfaces spanning tens of metres, and severe class imbalance. We introduce MonuSegFormer, a heritage-specific hybrid architecture coupling a pretrained Swin-B encoder (ImageNet-22K) with a Multi-Scale Atrous Fusion (MSAF) module and a CBAM-augmented progressive decoder. Evaluated on the Moroccan Monuments Dataset (MMD, 2686 annotated RGB images, 5 semantic classes, 5 cities) on the held-out test set (403 images), MonuSegFormer achieves mIoU = 81.7%, outperforming SegFormer-B5 (76.4%), Mask2Former (78.9%), and DeepLabV3+ (71.3%). A composite CE+Dice+Focal loss with class-balanced weights addresses severe class imbalance, yielding the largest gains on the minority classes Door (+5.3 p.p.) and Window (+5.3 p.p.) over the next-best Mask2Former. All qualitative results are produced by real trained-model inference on held-out test images.