Underwater short-utterance speaker identification via Active Identity Embedding and parameter-efficient adaptation
Tianhao Luo1,2,3, Feng Zhou1,2,3,4, Gang Qiao1,2,3,4
1National Key Laboratory of Underwater Acoustic Technology, Harbin Engineering University, Harbin 150001, China.
Abstract:
Ensuring diver safety during emergency scenarios requires robust speaker identification systems. However, traditional passive voiceprint extraction faces two major limitations: (1) distress utterances are often extremely short and (2) severe underwater channel distortions degrade the acoustic-phonetic features. To break these limitations, we propose a paradigm shift toward Active Identity Embedding, which reformulates speaker identification as an active message embedding and extraction task. Our approach integrates identity embedding and extraction modules into the HiFi-Codec architecture, utilizing a multi-dilation Gated Convolutional Unit-based decoder to resist significant inter-symbol interference. To facilitate model adaptation in dynamic underwater environments, we propose a two-stage pretraining-fine-tuning strategy. This strategy employs a hybrid fine-tuning scheme-combining low-rank adaptation, full-parameter fine-tuning (FPFT), and parameter freezing-to minimize the computational overhead for on-device updates whereas acting as a structural regularizer. Experimental results on LibriTTS and WATERMARK datasets show that the proposed method successfully embeds a 16-bit identity message into 1-s speech segments with minimal quality degradation (perceptual evaluation of speech quality narrowband, PESQ-NB ≤ 0.07), achieving a bit error rate < 1% at 15 dB signal-to-noise ratio. Furthermore, the hybrid fine-tuning scheme reduces trainable parameters to only 18.7% of those required for FPFT.
Related Concept Videos
Impression Management Techniques IV: Altercasting
Nonconscious Mimicry
Social Identity
Personal Identity
Role-Based Identity
