基本概率模型用于文本同质性和细分的应用:圣经的一个案例研究
1Department of Statistics, MaiNefhi College of Science, Mai-Nefhi, Zoba Maekel, Eritrea.
PloS one
|June 7, 2024
概括
对提格里纳语,阿姆哈拉语和英语的圣经翻译进行统计分析,揭示了文本异质性. 概率模型显示,圣经的翻译是混合的类型,保罗的信包括两个不同的部分.
科学领域:
- 计算语言学 计算语言学
- 统计分析 统计分析
- 文学研究 文学研究
背景情况:
- 该研究应用了新的概率模型来分析文本特征.
- 圣经文本是复杂的,包括各种类型和文学风格.
- 了解文本的一致性对于准确的解释和翻译分析至关重要.
研究的目的:
- 用先进的概率模型统计测试圣经书籍.
- 分析文本的一致性,并检测圣经翻译中的变化点.
- 调查提格里纳,阿姆哈拉语和英语圣经翻译的结构组成.
主要方法:
- 利用最近开发的概率模型用于文本同质性和变化点检测.
- 运用统计测试对Tigrigna,阿姆哈拉语和英语的圣经翻译.
- 分析了文本细分和分布模式,包括Zipf-Mandelbrot分布.
主要成果:
- 在所有三种圣经翻译中都发现了Zipf-Mandelbrot分布 (参数范围为0.55-0.88).
- 统计分析表明,每个圣经翻译都是不同书籍或不同类型的异构混合物.
- 对英语保罗的信件进行了深入的细分,发现它们由两个同质的细分组成.
结论:
- 提格里纳语,阿姆哈拉语和英语的圣经翻译表现出相当大的文本异质性.
- 这些发现表明,圣经的翻译不是单一的,而是不同文本单元的连接.
- 特别是保罗的书信,展示了由两个单独的同质部分组成的复杂的内部结构.
更多相关视频
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.2K
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
8.7K
相关概念视频
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Probability in Statistics
12.7K
Probability is the likelihood of an event occurring. The term event is defined as a collection of results of a procedure. An event is a simple event when an outcome cannot be divided into simpler parts.
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
12.7K
Poisson Probability Distribution
7.8K
A Poisson probability distribution is a discrete probability distribution. It gives the probability of a number of events occurring in a fixed interval of time or space if these events happen at a known average rate and independently of the time since the last event. For example, a book editor might be interested in the number of words spelled incorrectly in a particular book. It might be that, on average, there are five words spelled incorrectly in 100 pages. The interval is 100 pages.
The...
The...
7.8K
Probability Distributions
6.9K
The probability of a random variable x is the likelihood of its occurrence. A probability distribution represents the probabilities of a random variable using a formula, graph, or table. There are two types of probability distribution– discrete probability distribution and continuous probability distribution.
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
6.9K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Stratified Sampling Method
12.0K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
12.0K
