关于最小化器和卷积过器:对基因组分析的理论联系和应用
1Department of Mathematics, University of Toronto, Toronto, Ontario, Canada.
概括
随机高斯初始化和最大聚合的卷积神经网络 (CNN) 在数学上相当于基于最小化器的分析生物序列的方法. 这种等价性解释了它们的有效性,特别是在重复的基因组区域.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 缩小器和卷积神经网络 (CNN) 是分析生物序列的不同技术.
- 最小化器使用k-mer散列,而CNN则使用过器和聚合器进行特征提取和分类.
研究的目的:
- 以数学分析连接最小化器和CNN的哈希函数属性.
- 解释CNN在分类序列分析中的有效性.
主要方法:
- 哈希函数属性的数学分析.
- 在模拟和真实生物序列 (人类端粒) 上进行实证实验.
- 训练一个CNN嵌入SARS-CoV-2基因组短读.
主要成果:
- 随机高斯初始化CNN过器与最大共享相当于特定的最小化订单.
- 这种等价性导致重复序列区域的k-mer密度下降.
- 一个CNN嵌入的SARS-CoV-2在本地读取回合序列距离,虽然目前不切实际.
结论:
- 为生物序列分析中的CNN有效性提供了部分数学解释.
- 突出了看似不相似的序列分析技术之间的联系.
- 暗示了在序列组装中深度学习的潜力.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Genomics
36.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.3K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K


