对古代基因组的基准测试亲属关系估计工具使用血统模拟.
Şevval Aktürk1, Igor Mapelli1, Merve N Güler1
1Department of Biological Sciences, Middle East Technical University, Ankara, Turkey.
Molecular ecology resources
|April 27, 2024
概括
研究人员评估了从低覆盖率的古代DNA (古遗传组) 估计遗传亲属关系的四种工具. NgsRelate和lcMLkin显示出高精度的近亲,即使有限的单核酸多态 (SNP).
科学领域:
- 古代的基因组学 古代的基因组学
- 人口遗传学 人口遗传学
- 生物信息学是一种生物信息学.
背景情况:
- 越来越多的人对使用低覆盖范围的古遗传基因组重建古代亲属关系模式感兴趣.
- 需要强大的亲属估计工具,适用于有限的古代遗传数据.
研究的目的:
- 在低覆盖范围的古遗传学数据上对四种亲属关系估计工具 (lcMLkin, NgsRelate, KIN, READ) 的性能进行基准和比较.
- 为了评估不同数量的单核酸多态 (SNPs) 的工具准确性,并评估人口等位基因频率噪声和近亲繁殖的影响.
主要方法:
- 利用了血统和古代基因组序列模拟,具有有限的共享SNP (1至50K,MAF≥0.01).
- 对不同程度的亲属关系 (一,二,三等亲属) 评估了亲属关系估计的准确性.
- 研究了人口等位基因频率噪声和近亲繁殖对亲属系数估计的影响.
主要成果:
- 所有四种工具的性能与≥20KSNP相比.
- 一级亲属被准确地分类为只有1KSNP (F1得分高达96%与NgsRelate/lcMLkin).
- 区分三级亲属的准确性很高 (F1>90%),使用NgsRelate和lcMLkin的5KSNP,表现优于READ和KIN.
结论:
- NgsRelate和lcMLkin在古遗传学数据中使用有限的SNP来证明近亲和远亲的高精度.
- 种群的等位基因频率噪音和近亲繁殖可以影响亲属系数,具有不同的工具灵敏度.
- 建议同时使用多个亲属关系估计工具,以获得超低覆盖率古老基因组的强大结果.
相关概念视频
Pedigree Analysis
84.2K
Overview
84.2K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Incomplete Dominance
22.5K
Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
22.5K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Next-generation Sequencing
88.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.7K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K


