盖亚:一个支持人工智能的基因组上下文意识平台,用于蛋白质序列注释
Nishant Jha1, Joshua Kravitz1, Jacob West-Roberts1
1Tatta Bio, Cambridge, MA 02142, USA.
Science advances
|June 20, 2025
概括
盖亚 (基因组人工智能解读器) 是一种用于蛋白质序列相似性搜索的新工具. 它使用基因组背景来找到其他方法遗漏的功能相关基因,改进生物研究.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 蛋白质序列相似性搜索在生物学中至关重要.
- 当前的方法往往忽视了基因组背景,限制了功能洞察力,特别是在微生物中.
- 在保存的基因组环境中识别功能相关的基因是具有挑战性的.
研究的目的:
- 介绍Gaia (基因人工智能解读器),这是一个用于快速,上下文感知蛋白质序列搜索的新平台.
- 为了利用基因组背景来改进功能相关基因的识别.
- 为微生物基因组分析提供免费可用的基于网络的工具.
主要方法:
- 开发了Gaia,一个使用混合模式基因组语言模型 (gLM2) 的序列注释平台.
- 训练了gLM2对氨基酸序列及其基因组邻居进行训练,以产生集成的嵌入.
- 实现了来自131,744个微生物基因组的8500多万个蛋白质的数据库中的实时搜索.
主要成果:
- 盖亚的嵌入整合了序列,结构和基因组上下文信息.
- 在保存的基因组环境中识别了功能和/或进化相关的基因.
- 与传统的嵌入和基于对齐的方法相比,证明了优越的同类检索性能.
结论:
- 盖亚通过结合关键的基因组背景来增强蛋白质序列搜索.
- 该平台有助于发现传统方法可能错过的基因关系.
- 盖亚为微生物基因组学研究提供了一个有价值的,可访问的资源.
相关概念视频
Genome Annotation and Assembly
19.4K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.4K
Genomics
37.5K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
37.5K
RNA-seq
10.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.4K
Next-generation Sequencing
92.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
92.8K
Cis-regulatory Sequences
3.1K
3.1K
Conservation of Protein Domains Over Different Proteins
11.5K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.5K


