CGRclust:混沌游戏表示用于未标记的DNA序列的双重对比集群
Fatemeh Alipour1, Kathleen A Hill2, Lila Kari3
1School of Computer Science, University of Waterloo, Waterloo, Canada. falipour@uwaterloo.ca.
BMC genomics
|December 19, 2024
概括
本研究介绍了CGRclust,这是一种使用混乱游戏表示 (CGR) 和卷积神经网络 (CNN) 的新型无监督DNA序列聚类方法. CGRclust提供了准确,可扩展和无对齐的DNA序列分类,没有标签,优于现有的方法,特别是病毒基因组.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 传统的DNA序列分类依赖于劳动密集型标记和耗时的序列对齐.
- 这些局限性阻碍了大型基因组数据集的可扩展性和远距离相关生物体的分析.
- 需要有效的,无监督的DNA序列聚类方法来绕过标记和对齐.
研究的目的:
- 推出CGRclust,一种新的无监督DNA序列聚类方法.
- 通过使用无监督学习来克服传统方法的局限性,用于DNA序列分析.
- 为DNA序列聚类提供一种强大,可扩展和无对齐的方法.
主要方法:
- 使用混沌游戏表示 (CGR) 将DNA序列转换为图像.
- 应用无监督双胞胎对比学习与卷积神经网络 (CNN) 进行集群CGR图像.
- 在不需要序列对齐或分类标签的情况下检测出独特的序列模式.
主要成果:
- CGRclust精确地聚集了不同长度的25个不同的DNA序列数据集 (664bp到100kbp).
- 与DeLUCS,iDeLUCS和MeShClust (v3.0.0.3) 相比,鱼类线粒体DNA在所有分类学层次上都表现出卓越的准确性.
- 在病毒全基因组组装数据集上持续优于竞争方法,显示出高强度和多功能性.
结论:
- CGRclust是一种新的,可扩展和无对齐的DNA序列聚类方法.
- 与当前方法相比,实现更高或可比的准确性和性能,特别是对于病毒DNA数据集.
- 在90%以上的分析数据集中显示出超过80%的准确性,提高了可靠性.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Labeling DNA Probes
8.1K
DNA probes are fragments of DNA labeled with a reporter tag to enable their detection or purification. The resulting labeled DNA probes can then hybridize to target nucleic acid sequences through complementary base-pairing, and may be used to recover or identify these regions.
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
8.1K
DNA as a Genetic Template
21.7K
Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
21.7K
RNA-seq
9.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.8K
DNA Microarrays
17.2K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.2K
Karyotyping
57.5K
Overview
57.5K


