Related Experiment Video
Updated: Jul 8, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Speech recognition datasets for low-resource Congolese languages
Ussen Kimanuka1, Ciira Wa Maina2,3, Osman Büyük4
1Department of Electrical Engineering, Pan African University Institute for Basic Sciences, Technology and Innovation, Nairobi, Kenya.
New speech recognition datasets for Lingala and other Congolese languages were created. These resources aid in developing advanced Automatic Speech Recognition (ASR) models for low-resource languages.
Area of Science:
- Computational Linguistics
- Speech Processing
- Low-Resource Language Technologies
Background:
- Large pre-trained Automatic Speech Recognition (ASR) models benefit from transfer learning, but require substantial data.
- Many low-resource languages lack sufficient data to fully leverage transfer learning.
- Benchmark corpora are essential for advancing ASR methods in data-scarce linguistic contexts.
Purpose of the Study:
- To introduce two novel benchmark corpora for low-resource languages in the Democratic Republic of the Congo.
- To facilitate the development of monolingual and multilingual ASR systems for Congolese languages.
- To enable inaugural benchmarking of speech recognition systems for Lingala and four other Congolese languages.
Main Methods:
- Creation of the Lingala Read Speech Corpus (4h labeled audio) with diverse speakers and accents.
- Compilation of the Congolese Speech Radio Corpus (741h unlabeled audio) from broadcast archives.
- Application of supervised learning and self-supervised learning techniques for model development and benchmarking.
Main Results:
- The developed corpora provide valuable resources for ASR research in low-resource settings.
- Successful inaugural benchmarking of speech recognition systems for Lingala.
- Development of the first multilingual ASR model for four Congolese languages, serving 95 million people.
Conclusions:
- The released datasets are crucial for advancing ASR research and development for underserved languages.
- These resources pave the way for improved speech recognition technologies in the Democratic Republic of the Congo.
- The study highlights the potential of transfer learning and novel corpora for low-resource ASR.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
Related Concept Videos
Sound Intensity Level
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Korotkoff Sounds
During blood pressure assessment, inflating the cuff 30 millimeters of mercury above the patient's systolic blood pressure...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Sound Intensity