Acoustic feature decoupling and pre-trained language model integration for music-to-text cross-modal generation

Jiaying Yu1

  • 1Graduate School of Hanyang University, Seoul, 04763, Korea. yujiaying25@163.com.

Scientific Reports
|July 1, 2026
PubMed
Summary

This study introduces a new framework for music captioning, effectively translating audio into text by decoupling musical features and using advanced language models. The approach significantly improves the quality and relevance of generated music descriptions.

Related Concept Videos