链接用于准确对准错误的长读数到非循环变化图的错误长读数
Jun Ma1, Manuel Cáceres1, Leena Salmela1
1Department of Computer Science, University of Helsinki, 00014 Helsinki, Finland.
Bioinformatics (Oxford, England)
|July 26, 2023
概括
GraphChainer通过同线性链接多个种子来改善长读对齐到泛基因组学变异图. 与现有方法相比,这种新的对齐器显著增加了对齐读数的数量和长度.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 将序列阅读与变异图对齐在泛基因组学中对于变异调用等应用至关重要.
- 像vg这样的现有工具在短时间读取中很受欢迎,而GraphAligner则是长时间错误读取的最先进工具.
- GraphAligner的种子扩展方法可以通过共同线性链接多个种子来改进.
研究的目的:
- 介绍 GraphChainer,这是一个新的算法和对齐器,用于在非循环变化图中同线性链接种子.
- 有效地将错误的长读数与复杂的泛基因组变异图进行对齐.
主要方法:
- 开发了一种新的算法,用于在字符串标记的非循环图中对种子进行同线性链接.
- 在 GraphChainer 调整器中实现了一种高效的同线性链接算法.
- 在真实和模拟的PacBio CLR上评估GraphChainer的读数具有不同的错误率.
主要成果:
- 在人类泛基因组数据上,GraphChainer比GraphAligner对准了12-17%的阅读量和21-28%的总阅读长度.
- 在模拟和真实数据上实现了95-99%的读取和总读取长度对齐.
- 与微图和微链相比,表现出更高的性能,其精度为<60%.
结论:
- GraphChainer代表了长读对应变量图的显著进步.
- 同线性链接方法提高了对齐的准确性和完整性.
- GraphChainer为泛基因组分析提供了一个更强大,更有效的解决方案.
相关概念视频
Genome Copying Errors
4.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.3K
Mismatch Repair
40.3K
Overview
40.3K
Proofreading
54.2K
Overview
54.2K
Sanger Sequencing
754.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.8K
Comparing Copy Number Variations and SNPs
17.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.8K
Improving Translational Accuracy
11.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.6K


