在杂的基因组序列中检测基因断点,使用位置注释的彩色de-Bruijn图表.
Lisa Fiedler1, Matthias Bernt2, Martin Middendorf3
1Department of Computer Science, University Leipzig, Augustusplatz 10-11, 04109, Leipzig, Germany. lfiedler@informatik.uni-leipzig.de.
BMC bioinformatics
|June 5, 2023
概括
这项研究引入了DeBBI,这是一种用于在线粒体DNA序列中找到基因断点的新方法,即使变化很高. DeBBI准确地检测进化中断点,改善了线粒体基因组进化的分析.
科学领域:
- 基因组学就是基因组学.
- 进化生物学 进化生物学
- 生物信息学是一种生物信息学.
背景情况:
- 识别基因断点有助于理解进化过程.
- 线粒体基因组由于基因顺序和序列不一致的高变化而存在挑战.
- 在核酸序列中准确检测断点是很困难的错误注释或缺失的基因位置.
研究的目的:
- 提出一种用于检测线粒体基因组核酸序列中的基因断点的新方法.
- 开发一个软件包DeBBI,用于分析基于转换和反转的断点.
- 为应对高替代率和序列不一致性所带来的挑战.
主要方法:
- 开发了DeBBI软件包,使用了一种新的断点检测方法.
- 采用并行程序设计,用于对多处理器系统的高效分析.
- 构建了一个位置注释的de-Bruijn图,并使用启发式算法找到与断点相关的结构 (凸起).
主要成果:
- DeBBI在合成数据集上展示了准确的结果,其中有不同的序列差异和断点数.
- 案例研究证实了DeBBI对各种分类群体的真实数据的适用性.
- 与一些对齐工具相比,DeBBI显示了改进的基因断裂检测,特别是在短,保存不良的tRNA基因之间.
结论:
- 拟议的方法有效地识别了线粒体核酸序列中的基因断点.
- DeBBI提供了一种强大而准确的方法来分析线粒体基因组进化.
- 德布鲁因图和启发式算法提供了一个有效的方法来定位断点.
相关概念视频
Gene Duplication and Divergence
6.2K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.2K
Comparing Copy Number Variations and SNPs
17.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.8K
Genome Copying Errors
4.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.3K
Evolutionary Relationships through Genome Comparisons
5.9K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.9K
Sanger Sequencing
755.2K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
755.2K
DNA Microarrays
17.6K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.6K


