Related Experiment Video
Updated: Jan 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Enhancing end-to-end speech translation via multi-stage knowledge distillation
Yue Zhou1, Yuxuan Yuan1, Yanyan Feng2
1School of Informatics, Xiamen University, Xiamen, 361005, Fujian, China; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan, Ministry of Culture and Tourism, China.
This study introduces multi-grained knowledge distillation for speech-to-text translation, improving knowledge transfer from teacher models. The new method enhances translation quality by focusing on difficult translations and aligning cross-modal representations.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Machine Learning
Background:
- Knowledge distillation (KD) enhances speech-to-text translation (ST) using machine translation (MT) teacher models.
- Current KD methods transfer limited knowledge by relying on teacher output distributions, hindering student model performance.
- Scarcity of ST data further limits comprehensive knowledge transfer from MT to ST models.
Purpose of the Study:
- To propose a multi-grained distillation method for more effective knowledge transfer in ST.
- To address limitations of existing KD methods in capturing deep representations and handling data scarcity.
- To improve the overall quality and effectiveness of end-to-end speech-to-text translation.
Main Methods:
- Introduced adaptive word-level distillation to prioritize challenging translations.
- Implemented cross-modal hidden state distillation to align MT and ST model representations, bridging the speech-text modality gap.
- Developed a multi-stage knowledge distillation (MSKD) framework leveraging external ASR, MT, and ST data.
Main Results:
- MSKD achieved state-of-the-art performance on the MuST-C dataset, outperforming previous KD methods by +2.5 BLEU.
- With external data, MSKD surpassed strong end-to-end ST baselines by +2.9 BLEU and cascaded systems by +1.9 BLEU.
- Demonstrated significant improvements in ST performance, highlighting the method's effectiveness and scalability.
Conclusions:
- The proposed multi-grained distillation method significantly enhances knowledge transfer for ST.
- MSKD effectively utilizes diverse data sources and distillation techniques for progressive ST performance improvement.
- The framework offers a scalable and effective approach to advancing end-to-end speech-to-text translation capabilities.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Impression Management Techniques IV: Altercasting
Translation
Translation Produces the Building Blocks of Life
Proteins are...
Translation
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of...
Termination of Translation
