Related Experiment Video
Updated: Jun 12, 2025

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
Machine translation training data for English-Tshivenḓa
Tanja Gaustad1, Cindy A McKellar1, Martin J Puttkammer1
1Centre for Text Technology (CTexT), North-West University, 11 Hoffman Street, Potchefstroom, South Africa.
Abstract:
This data article describes a machine translation training data set for translation between English and Tshivenḓa. The data set contains parallel, aligned English-Tshivenḓa data as well as monolingual Tshivenḓa data. The data was collected from both web crawling of multilingual South African government sites and matched documents from translators or publishing sources. Additional unique data was translated from English into Tshivenḓa by professional translators to increase the total corpus size. This article contains information about the collection and translation of the data as well as how alignments and corpus cleanup were done. The wordcounts of the corpus are also given. In addition to training machine translation systems this data can also be used for the development of other Tshivenḓa core technologies as well as for linguistic studies.
More Related Videos
Related Concept Videos
Translation
Translation Produces the Building Blocks of Life
Proteins are...
Improving Translational Accuracy
Initiation of Translation
Termination of Translation
Transfer RNA Synthesis
Each of these chemical modifications is carried by a specific enzyme, post-transcription. All of these enzymes have unique base and site-specificity. Methylation, the most common chemical modification, is carried by at least nine different enzymes, with...
tRNA Activation

