核密度估计的等位基频率,包括未检测到的等位基
Satoshi Aoki1, Keita Fukasawa1
1Biodiversity Division, National Institute for Environmental Studies, Tsukuba, Ibaraki, Japan.
PeerJ
|April 26, 2024
概括
核密度估计 (KDE) 被探索用于估计使用未检测到的等位基因的遗传多样性. 然而,这种方法在等位基因频率和核酸多样性方面显示出更差的偏差和准确性,这表明它对遗传多样性估计没有好处.
科学领域:
- 人口遗传学 人口遗传学
- 生物信息学是一种生物信息学.
- 统计生态学统计生态学
背景情况:
- 未被发现的物种影响物种多样性的估计.
- 尚未检测到的等位基因并未用于遗传多样性估计.
- 虽然随机抽样提供了公正的估计,但未被检测到的等位基因可以提供有偏见但更精确的保护见解.
研究的目的:
- 开发和评估基因密度估计 (KDE) 的等位基因频率,包括未检测到的等位基因.
- 评估KDE在估计等位基因频率和核酸多样性的实用性.
主要方法:
- 核密度估计 (KDE) 应用于等位基频率数据.
- 该方法使用凝聚模拟和真实人口遗传数据进行了测试.
主要成果:
- 与预期相反,KDE导致核酸多样性估计的偏差增加和精度降低.
- 基于KDE的等位基频率估计通常更差,除了非常小的样本大小的情况下.
- 性能差的潜在原因包括人口有限性和维度的诅咒.
结论:
- 对等位基因频率的核密度估计 (KDE) 并不能改善遗传多样性的估计.
- 通过KDE将未被检测的等位基因纳入的方法目前对人口遗传分析没有好处.
相关概念视频
Hardy-Weinberg Principle
72.1K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
72.1K
What is Population Genetics?
57.9K
A population is composed of members of the same species that simultaneously live and interact in the same area. When individuals in a population breed, they pass down their genes to their offspring. Many of these genes are polymorphic, meaning that they occur in multiple variants. Such variations of a gene are referred to as alleles. The collective set of all the alleles within a population is known as the gene pool.
57.9K
Mutation, Gene Flow, and Genetic Drift
58.4K
In a population that is not at Hardy-Weinberg equilibrium, the frequency of alleles changes over time. Therefore, any deviations from the five conditions of Hardy-Weinberg equilibrium can alter the genetic variation of a given population. Conditions that change the genetic variability of a population include mutations, natural selection, non-random mating, gene flow, and genetic drift (small population size).
58.4K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Genetic Drift
39.7K
Natural selection—probably the most well-known evolutionary mechanism—increases the prevalence of traits that enhance survival and reproduction. However, evolution does not merely propagate favorable traits, nor does it always benefit populations.
39.7K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K


