从蛋白质语言模型中通过局部对准小位置嵌入的敏感远程同质性搜索
Sean R Johnson1, Meghana Peshwa1, Zhiyi Sun1
1New England Biolabs Inc, Ipswich, United States.
eLife
|March 15, 2024
概括
深度学习嵌入显著改善蛋白质进化关系的检测. 这些低维度的位置嵌入增强了对同质性搜索的灵敏度,而不会减缓该过程.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习在基因组学中的应用
背景情况:
- 检测遥远的蛋白质进化关系是具有挑战性的,特别是在<20%的序列相同性.
- 目前基于配置文件和结构的方法是敏感的,但由于缓慢的预处理,计算密集.
- 深度学习嵌入式提供更快的替代方案,但在区分域同质性和搜索速度方面面临限制.
研究的目的:
- 在敏感和快速的蛋白质同质性搜索中研究低维度位置嵌入的实用性.
- 将深度学习嵌入式集成到现有的优化搜索算法中,以改善进化分析.
主要方法:
- 利用ESM2 3B模型从蛋白质初级序列生成定位嵌入.
- 转换嵌入到3D交互 (3Di) 字母和氨基酸配置文件中的嵌入.
- 集成这些嵌入为输入优化搜索工具:Foldseek,HMMER3和HH-suite.
主要成果:
- 证明了低维度的位置嵌入 (仅为一个字节) 可以捕获足够的信息来检测同质性.
- 与标准的氨基酸序列搜索相比,实现了显著提高的灵敏度.
- 保持或改进了搜索速度,克服了以前基于嵌入的方法的局限性.
结论:
- 低维度的定位嵌入对于在很长的进化距离上敏感的蛋白质同质检测是有效的.
- 将嵌入式直接集成到速度优化的本地搜索算法中是可行的和高效的.
- 这种方法增强了进化关系推断,而不会影响计算性能.
相关概念视频
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Conservation of Protein Domains
3.1K
3.1K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Nuclear Localization Signals and Import
5.7K
Proteins targeted to the nucleus carry short stretches of amino acid sequences called the nuclear localization signal or NLS. Classical nuclear localization signals are of two types: monopartite and bipartite NLS. Monopartite classical NLS (cNLS) consists of a single cluster of 4-8 amino acids. Bipartite cNLS consists of two clusters of 2-3 amino acids and a 9-12 residue long proline-rich linker bridging the two clusters. Signal clusters are rich in positively charged amino acids such as...
5.7K


