Related Experiment Videos
Spectrum-Adaptive Modulation for Generalizable RGB-to-Thermal Semantic Segmentation
Abstract:
Robust scene understanding in adverse conditions typically relies on thermal imagery, as RGB sensors often fail under low illumination, fog, or smoke. This leads to an RGB-to-thermal semantic segmentation setting, where models trained on abundant labeled RGB images must generalize to thermal images at test time. Unlike conventional domain shifts, this setting involves a fundamental cross-spectral discrepancy: RGB and thermal sensors operate in distinct wavelength ranges and exhibit markedly different frequency characteristics. Although recent approaches leverage vision foundation models (VFMs) with feature alignment or augmentation strategies, they struggle to overcome the substantial spectral discrepancy and the absence of thermal-specific priors. To address this problem, we exploit spectral decomposition to jointly capture modality-specific patterns and modality-invariant semantic structures. We propose Spectrum-Adaptive Modulation (SAMO), which uses a small set of unlabeled thermal reference images to guide the modulation of RGB features toward the thermal domain during training. SAMO decomposes RGB and thermal inputs into low-, mid-, and high-frequency components, selectively preserving semantically relevant RGB features while recomposing thermal-specific components to mitigate the cross-spectral discrepancy. A mutual-information-driven objective and a spectrum-adaptive LoRA module are further proposed to enhance spectral decomposition and enable efficient VFMs adaptation. For comprehensive evaluation, we establish the RGB-to-Thermal Semantic Segmentation (RTSS) benchmark based on widely used driving-scene datasets. Extensive experiments show that SAMO achieves state-of-the-art performance, outperforming recent domain generalization methods. The dataset and code are available at: https://github.com/RTZhang98/SAMO-for-RGB2T.