Related Experiment Video
Updated: Jul 2, 2026

09:09
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Acoustic feature decoupling and pre-trained language model integration for music-to-text cross-modal generation
1Graduate School of Hanyang University, Seoul, 04763, Korea. yujiaying25@163.com.
Scientific Reports
|July 1, 2026
Summary
This study introduces a new framework for music captioning, effectively translating audio into text by decoupling musical features and using advanced language models. The approach significantly improves the quality and relevance of generated music descriptions.
Area of Science:
- Artificial Intelligence
- Music Information Retrieval
- Natural Language Processing
Background:
- Generating natural language descriptions from music audio is challenging due to complex acoustic representations and the semantic gap between audio and text.
- Existing methods struggle to effectively disentangle and represent diverse musical attributes like content, style, and emotion.
Purpose of the Study:
- To develop a novel framework for music captioning that integrates acoustic feature decoupling with large-scale pre-trained language models.
- To improve the coherence, relevance, and informativeness of automatically generated music descriptions.
Main Methods:
- A variational autoencoder (VAE) module factorizes audio into content, style, and emotion subspaces using mutual information minimization and orthogonality constraints.
- A multi-granularity alignment mechanism connects decoupled acoustic features with linguistic representations via contrastive learning.
- A pre-trained language model decoder is adapted using prefix tuning for multimodal conditioning.
Main Results:
- The proposed framework achieves competitive performance on three benchmarks (MusicCaps, Song Describer, LP-MusicCaps).
- Demonstrated relative gains of 18.9% in BLEU-4 and 4.4% in BERTScore over the CLAP-Cap baseline on MusicCaps.
- Ablation studies and human evaluations confirm the contribution of each component and the superior quality of generated descriptions in fluency, relevance, informativeness, and diversity.
Conclusions:
- The framework effectively bridges the modality gap in music captioning by decoupling acoustic features and leveraging pre-trained language models.
- The proposed approach offers a significant advancement in generating high-quality, semantically rich descriptions for music audio.
- This work paves the way for more sophisticated cross-modal generation tasks involving complex audio data.
