相关实验视频
Updated: Jun 29, 2025

08:57
Using Phylogenetic Analysis to Investigate Eukaryotic Gene Origin
Published on: August 14, 2018
15.9K
准确地通过相关性分类在线时间中对生物序列进行聚类
Erik Wright1,2
1Department of Biomedical Informatics, University of Pittsburgh, Pittsburgh, PA, USA. eswright@pitt.edu.
Nature communications
|April 8, 2024
概括
一个新的算法,Clusterize,有效地以高精度将数百万个生物序列聚集在一起. 这种方法实现了线性时间复杂性,优于现有的线性时间算法,并与较慢,更准确的用于大规模生物数据分析的算法竞争.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 由于数据的指数增长,生物序列的聚类至关重要.
- 对于大型数据集而言,现有的超线性聚类方法在计算上昂贵.
- 当前的线性时间算法为了速度而牺牲了准确性.
研究的目的:
- 开发一个具有高精度的线性时间序列聚类算法.
- 描述新算法的性能.
- 能够有效地对数以百万计的生物序列进行聚类.
主要方法:
- 开发了Clusterize,这是一种根据相关性对序列进行排序的算法.
- 通过序列排序线性化了聚类问题.
- 评估Clusterize与已建立的工具 (如CD-HIT,MMseqs2,UCLUST和Linclust) 相比.
主要成果:
- 集群实现准确度与流行的超线性方法相提并论.
- 证明了线性异常可扩展性,对于大数据集来说是实用的.
- 在线时间集群的精度和集群大小上都优于Linclust.
- 成功聚集了数百万个核酸和蛋白质序列.
结论:
- 聚类为生物序列聚类提供了一个可扩展和准确的解决方案.
- 为分析大规模的基因组和蛋白质组数据提供了有价值的工具.
- 解决了在生物信息学中对高效而又精确的序列分组的需求.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Phylogenetic Trees
45.3K
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.
45.3K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
RNA-seq
9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Next-generation Sequencing
88.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.7K
Signal Sequences and Sorting Receptors
5.4K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.4K

