Related Experiment Video
Updated: Mar 19, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Cross-modal collaborative optimization for semantic segmentation of RGB-T power equipment images
Abstract:
Semantic segmentation is essential for detailed scene interpretation in complex substation environments. Leveraging the complementary properties of RGB and thermal infrared imaging, RGB-T multimodal data can enhance feature representation in semantic tasks. However, most existing approaches rely on single-stage fusion strategies, which overlook the modality-specific advantages at different semantic levels. In this work, we propose CCONet, a cross-modal collaborative optimization network designed for RGB-T semantic segmentation of power equipment. CCONet employs a three-stage framework-comprising feature extraction, multimodal fusion, and decoding-with joint cross-modal interactions throughout. The model integrates a feature pyramid network and a path aggregation network to ensure hierarchical information flow. Furthermore, we introduce three task-specific fusion modules: a feature excitation module for high-level semantic enhancement, a feature localization module for mid-level spatial alignment, and a feature refinement module for improving low-level texture representation. Experimental results on a self-constructed RGB-T substation dataset and public benchmarks demonstrate that CCONet achieves competitive performance compared to state-of-the-art methods, highlighting its potential in robust multimodal image analysis for power system monitoring and inspection; moreover, although CCONet is designed and fine-tuned for semantic segmentation in complex substation environments, its cross-modal collaborative optimization strategy is general and can be extended to other scenarios with similarly complex conditions, given sufficient domain-specific retraining.
