西斯瓦提语数据集:英语和西斯瓦提语的平行文本数据和西斯瓦提语的单语言文本数据
Tanja Gaustad1, Cindy A McKellar1, Martin J Puttkammer1
1Centre for Text Technology, North-West University, South Africa.
Data in brief
|April 15, 2024
概括
这篇数据文章介绍了用于机器翻译和自然语言处理 (NLP) 开发的新英语-西斯瓦蒂语数据集. 该集体支持Autshumato机器翻译项目和Siswati的语言研究.
科学领域:
- 计算语言学 计算语言学
- 非洲语言 非洲语言
- 数据科学数据科学数据科学
背景情况:
- 西斯瓦蒂语是南非和埃斯瓦蒂尼的官方语言.
- 西斯瓦蒂的数字资源有限,阻碍了NLP的发展.
- 机器翻译和NLP技术需要大量的语言数据.
研究的目的:
- 为西斯瓦蒂语提供一个新的数据集.
- 为了支持Autshumato机器翻译项目.
- 为了促进自然语言处理 (NLP) 研究和西斯瓦蒂语的语料库语言学.
主要方法:
- 收集并行的英语-西斯瓦蒂文本数据.
- 编译的单语言西斯瓦蒂数据.
- 为NLP应用程序进行数据清理和预处理.
主要成果:
- 一个包括平行和单语言西斯瓦蒂文本在内的综合数据集.
- 详细描述数据收集和清理程序.
- 从词数来看,数据集的大小的概述.
结论:
- 这一数据集是推动西斯瓦蒂NLP和机器翻译的宝贵资源.
- 这些数据可用于开发和评估NLP工具.
- 该语料库还作为西斯瓦蒂语的语料库语言研究的基础.
相关概念视频
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Translation
14.8K
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Proteins are...
Translation Produces the Building Blocks of Life
Proteins are...
14.8K
Longitudinal Research
12.0K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
12.0K
Improving Translational Accuracy
2.6K
2.6K
Sanger Sequencing
754.2K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.2K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K


