Related Experiment Video
Updated: Nov 7, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.8K
Dynamic Acoustic Unit Augmentation with BPE-Dropout for Low-Resource End-to-End Speech Recognition
Aleksandr Laptev1, Andrei Andrusenko1, Ivan Podluzhny1
1Corporate Laboratory of Human-Machine Interaction Technologies, Information Technologies and Programming Faculty, School of Translational Information Technologies, ITMO University, 196084 Saint-Petersburg, Russia.
Sensors (Basel, Switzerland)
|April 30, 2021
Summary
This study introduces BPE-dropout for on-device automatic speech recognition (ASR) in low-resource settings. The technique improves recognition of unseen words and personalization without extra computational cost.
Area of Science:
- Speech Technology
- Machine Learning
- Natural Language Processing
Background:
- On-device automatic speech recognition (ASR) is crucial for speech assistants, with end-to-end models preferred for efficiency and quality.
- End-to-end ASR requires substantial speech data and struggles with personalization, particularly out-of-vocabulary (OOV) words.
- Low-resource scenarios with high OOV rates present significant challenges for developing effective ASR systems.
Purpose of the Study:
- To develop an effective end-to-end ASR system for low-resource conditions with high OOV rates.
- To address the challenge of recognizing OOV words in personalized ASR.
- To improve ASR performance without increasing computational demands.
Main Methods:
- Proposed a dynamic acoustic unit augmentation method using Byte Pair Encoding with dropout (BPE-dropout).
- BPE-dropout non-deterministically tokenizes utterances to enhance token context and regularize distributions.
- This method reduces the need for searching optimal subword vocabulary sizes.
Main Results:
- Achieved steady improvements in both regular and personalized (OOV-focused) ASR tasks.
- Demonstrated at least a 6% relative reduction in word error rate (WER) and a 25% relative increase in F-score.
- A monolingual Turkish Conformer model using BPE-dropout reached competitive results (22.2% CER, 38.9% WER).
Conclusions:
- BPE-dropout effectively enhances end-to-end ASR in low-resource, high-OOV environments.
- The technique offers significant performance gains without additional computational cost.
- The proposed method shows promise for improving ASR systems for personalized and resource-constrained applications.
Keywords:
BABEL GeorgianBABEL TurkishBPE-dropoutaugmentationend-to-end speech recognitionlow-resourceout-of-vocabularytransformerMore Related Videos
Related Concept Videos
Downsampling
347
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
347
Air-entraining Agents
150
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
150

