Related Experiment Video
Updated: May 8, 2026

Long-term Behavioral Tracking of Freely Swimming Weakly Electric Fish
Published on: March 6, 2014
Multi-scale self-distillation underwater acoustic signal recognition via saliency masking modeling
Xingmei Wang1,2, Zijian Liu1, Zheng Guo1
1College of Computer Science and Technology, Harbin Engineering University, Harbin 150001, China.
None:
A major challenge in underwater acoustic signal recognition is the lack of high-quality labeled data. Self-supervised learning leveraging masking modeling and reconstruction has emerged as an intuitive and feasible solution. The modeling process predicts and reconstructs the masked regions by learning the time-frequency features in the spectrogram. Despite the demonstrated success, these methods often suffer from overfitting to environmental noise and degraded recognition performance for weak targets, primarily due to neglecting cross-domain noise variations. To address this, we propose a multi-scale self-distillation framework via saliency masking modeling (SMM-MSSD) for underwater acoustic signal recognition. Saliency masking strategy takes spectrograms as input and masks patches with high attention scores in the self-attention map. SMM-MSSD predicts cross-domain invariant features by modeling contextual time-frequency information. Building on this, we design a multi-scale self-distillation architecture to align global and local information from weak targets at both the class and data levels. In addition, we introduce a pixel reconstruction loss to jointly optimize semantic perception and fine-grained feature modeling. Extensive experiments highlight the effectiveness of SMM-MSSD, achieving 88.52% and 82.36% accuracy on the ShipsEar and Mini-DeepShip datasets, respectively. Moreover, SMM-MSSD demonstrates strong robustness under low signal-to-noise ratios and limited training data.

