使用isONclust3对大型长时间读取的转录组数据集进行de novo聚类
Alexander J Petri1,2, Kristoffer Sahlin1
1Department of Mathematics, Science for Life Laboratory, Stockholm University, Stockholm 106 91, Sweden.
Bioinformatics (Oxford, England)
|April 23, 2025
概括
isONclust3有效地将大量长时间读取的转录组数据集成到基因家族中. 这种无参考方法比现有的工具快10-100倍,可以对大数据集和新生物进行分析.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 长读测序通过实现端到端分析来推进转录学研究.
- 基于参考的转录组分析工具面临着缺乏高质量的基因组或高度可变基因的生物的局限性.
- 长读数的无引用聚类对于分析这些具有挑战性的数据集至关重要.
研究的目的:
- 开发一种高效的算法,用于集群大规模的长读转录组数据集.
- 为了实现转录组的无引用分析,特别是对于具有有限基因组资源或高基因变异性的生物.
- 克服现有工具在处理大规模长读数据方面的计算局限性.
主要方法:
- 开发了isONclust3,这是一个改进的算法,用于将长读数聚类成基因家族.
- 采用了一种新的方法,使用动态集群表示与高可信度最小化器.
- 整合了一个代的集群合并步骤,以提高准确性和效率.
主要成果:
- isONclust3实现了比最先进的方法更高或可比的集群质量.
- 显示了显著的速度改进,在大型数据集上速度快10-100倍.
- 在单个计算节点上成功集群了3700万个PacBio读数,展示了高吞吐量测序的可扩展性.
结论:
- isONclust3提供了一个可扩展和高效的解决方案,用于无引用的长读转录组分析.
- 该算法克服了计算瓶,使各种生物系统的全面转录学研究成为可能.
- 简化了以前无法使用现有工具分析复杂的转录组的分析.
相关概念视频
RNA-seq
9.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.7K
Next-generation Sequencing
86.2K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
86.2K
Sanger Sequencing
751.3K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
751.3K
Genomics
35.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
35.3K
DNA Microarrays
17.1K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.1K
RACE - Rapid Amplification of cDNA Ends
6.2K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
6.2K


