Related Experiment Video
Updated: Jun 13, 2026

03:31
End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
Long-Tail Aware Cross-Modal Graph Attention Network for Fine-Grained Indoor 3D Semantic Segmentation of Point Clouds
Erdal Özbay1, Feyza Altunbey Özbay2
1Department of Computer Engineering, Firat University, Elazig 23119, Türkiye.
Sensors (Basel, Switzerland)
|June 12, 2026
Summary
This study introduces a novel network for 3D semantic segmentation, improving fine-grained object recognition in indoor scenes with imbalanced datasets. The Long-Tail Aware Cross-Modal Graph Attention Network (LT-CM-GACNet++) enhances rare class identification and overall segmentation accuracy.
Area of Science:
- Computer Vision
- 3D Data Analysis
- Machine Learning
Background:
- Accurate 3D semantic segmentation of indoor scenes is crucial for applications like robotics and augmented reality.
- High-resolution datasets like ScanNet++ present challenges due to fine-grained categories, high data density, and long-tail class distributions.
- Existing methods struggle with class imbalance and distinguishing visually/geometrically similar objects.
Purpose of the Study:
- To propose a novel network, LT-CM-GACNet++, for fine-grained 3D semantic segmentation addressing long-tail distributions.
- To effectively fuse geometric and RGB visual information for improved scene understanding.
- To enhance the learning of rare classes and improve discrimination between similar categories.
Main Methods:
- Developed a Long-Tail Aware Cross-Modal Graph Attention Network (LT-CM-GACNet++).
- Integrated dynamic graph-based geometric feature extraction with a lightweight MobileNetV3 visual feature extractor.
- Employed a Cross-Modal Graph Attention (CMGA) module for adaptive inter-modal information transfer.
- Utilized prototype-based representation learning and a class frequency-aware loss function to handle class imbalance.
- Applied preprocessing techniques including density-based sampling and normal vector estimation.
Main Results:
- The proposed LT-CM-GACNet++ achieved significant improvements in overall segmentation performance on the ScanNet++ dataset.
- Demonstrated notable gains in mean Intersection over Union (mIoU), particularly for rare classes.
- Outperformed existing approaches in both general segmentation accuracy and the challenging task of segmenting infrequent object categories.
Conclusions:
- Cross-modal learning is highly effective for high-resolution 3D scene segmentation, especially under long-tail data distributions.
- The proposed LT-CM-GACNet++ successfully addresses the challenges of fine-grained segmentation and class imbalance in indoor datasets.
- The method offers a promising direction for advancing 3D semantic segmentation in complex real-world scenarios.