Related Experiment Video
Updated: Oct 4, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
NCHLT Auxiliary speech data for ASR technology development in South Africa
Jaco Badenhorst1, Febe de Wet1,2
1Voice Computing Research Group, CSIR Next Generation Enterprises and Institutions Cluster, P.O. Box 395, Pretoria 0001, South Africa.
This study released new speech and text datasets for South Africa's official languages, aiding Human Language Technology (HLT) development. These auxiliary corpora help address the under-resourced nature of these languages for speech technology applications.
Area of Science:
- Computational Linguistics
- Speech Technology
- Corpus Linguistics
Background:
- South Africa has 11 official languages, all currently under-resourced for Human Language Technology (HLT) development.
- The National Centre for Human Language Technology (NCHLT) project aimed to create essential speech and text resources for these languages.
- Previous NCHLT releases, like the 2014 Speech Corpus, did not encompass all collected data.
Purpose of the Study:
- To describe and release additional speech and text data collected during the NCHLT project.
- To provide valuable resources for advancing speech technology for under-resourced South African languages.
- To supplement existing corpora and further enable HLT development.
Main Methods:
- Speech data collection via a smartphone application during the NCHLT project.
- Compilation and release of auxiliary speech corpora, including transcriptions for each utterance.
- Data curation and organization for accessibility and usability in HLT research.
Main Results:
- Release of auxiliary speech corpora containing 20 to 170 hours of data per language.
- Inclusion of transcriptions for all utterances within the auxiliary datasets.
- Significant expansion of available speech resources for South Africa's official languages.
Conclusions:
- The newly released auxiliary corpora significantly contribute to addressing the resource scarcity for HLT in South Africa.
- This data facilitates the development of speech technology for under-resourced languages.
- Further development of Human Language Technology for all 11 official South African languages is now more feasible.
Related Concept Videos
Air-entraining Agents
Automatic Processing and Automatic Social Behavior
RACE - Rapid Amplification of cDNA Ends

