Related Experiment Video
Updated: Jun 28, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Dataset for Siswati: Parallel textual data for English and Siswati and monolingual textual data for Siswati
Tanja Gaustad1, Cindy A McKellar1, Martin J Puttkammer1
1Centre for Text Technology, North-West University, South Africa.
Abstract:
This data article presents a dataset for Siswati, a Bantu language of the Nguni group that is one of the eleven official South African languages and the official language of Eswatini (together with English). The dataset contains parallel textual data between English and Siswati as well as monolingual data for Siswati and was developed for use as training data for machine translation systems, specifically the Autshumato machine translation project. Both corpora can also be used for development and evaluation of Natural Language Processing (NLP) core technologies for Siswati. In addition, the data lends itself for corpus linguistic studies. The article describes how the data was collected, what type of texts it contains and what clean-up was done. It also provides an overview of the number of words contained in the datasets.
Related Concept Videos
Test for Homogeneity
Translation
Translation Produces the Building Blocks of Life
Proteins are...
Longitudinal Research
Improving Translational Accuracy
Sanger Sequencing
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...

