DandD:高效测量序列增长和相似性
Jessica K Bonnie1, Omar Y Ahmed1, Ben Langmead1
1Department of Computer Science, Johns Hopkins University, Baltimore, MD, USA.
iScience
|February 16, 2024
概括
测量新的基因组序列内容是具有挑战性的. DandD使用新的三角形 (δ) 测量方法量化基因组组合中的序列增长,帮助泛基因组分析和预测发现和.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 基因组组装数据库正在迅速扩大.
- 在不断增长的集合中量化序列冗余性和新性是算法上复杂的.
- 现有的比较基因组组合的方法存在局限性.
研究的目的:
- 引入方法和工具 (DandD) 来测量在不断增长的序列集合中获得的新序列内容.
- 评估在人类基因组组合中发现结构变异的速度.
- 为了预测基因组组合中的新发现何时可能会平稳.
主要方法:
- 开发用于量化序列新性的DandD工具.
- 使用一种称为delta (δ) 的测量方法,该测量方法来自k-mer计数和基因组草图.
- 建议用 δ作为 k-mer 特定心的替代方案来计算 Jaccard 系数.
- 使用 k-独立的 Jaccard 进行相似性计算.
主要成果:
- DandD有效地估计了基因组组合中新序列内容的数量.
- 该工具描述了结构变异发现的速度.
- 在估计体生长率方面已证明有用.
- 启用了使用k独立的Jaccard有效的全对相似性计算.
结论:
- DandD提供了一种可靠的方法来测量大型基因组数据集中的序列新性.
- 德尔塔 (δ) 测量比传统的基于k-mer的比较提供了优势.
- 该工具有助于理解基因组组装完整性和泛基因组扩张的动态.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Per-Unit Sequence Models
74
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
74
Next-generation Sequencing
88.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.8K
Sanger Sequencing
754.3K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.3K
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K
Wald-Wolfowitz Runs Test I
647
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
647


