蛋白质语言模型是新的通用钥匙吗?
Konstantin Weissenow1, Burkhard Rost2
1TUM (Technical University of Munich), School of Computation, Information and Technology (CIT), Faculty of Informatics, Chair of Bioinformatics & Computational Biology - i12, Boltzmannstr. 3, 85748 Garching/Munich, Germany; TUM Graduate School, Center of Doctoral Studies in Informatics and its Applications (CeDoSIA), Boltzmannstr. 11, 85748 Garching, Germany.
Current opinion in structural biology
|February 8, 2025
概括
蛋白质语言模型 (pLMs) 为蛋白质预测提供了一种强大的新方法,在准确性和效率方面超过了传统方法. 这些模型使用嵌入式,以更少的计算资源提供蛋白质特定的见解.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 机器学习在生物学中的应用
背景情况:
- 传统的蛋白质预测在很大程度上依赖于来自多重序列对齐 (MSA) 的进化信息.
- 蛋白质语言模型 (pLMs) 学习蛋白质序列的固有语法.
- pLM嵌入式隐含地编码了这些学习的语法信息.
研究的目的:
- 评估pLM嵌入的有效性,作为蛋白质预测任务的独家输入.
- 将基于pLM的方法的性能和资源效率与传统的基于MSA的方法进行比较.
- 倡导优化现有的pLM基础模型,而不是重新培养新模型.
主要方法:
- 使用pLM生成的嵌入作为下游监督学习模型的唯一输入.
- 将预测准确度和计算资源消耗与已建立的基于MSA的方法进行比较.
- 评估pLM嵌入在凝结复杂蛋白序列信息中的效率.
主要成果:
- 无MSA的基于pLM的预测在许多应用程序中显示出显著提高的准确性.
- pLM嵌入式有效地缩小了序列语法,使得下游模型具有更少的参数.
- 基于pLM的解决方案提供蛋白质特定的预测,并且在预训后需要更少的计算资源.
结论:
- pLM正在迅速成为蛋白质预测的通用和高效工具.
- 基于pLM的方法比传统的基于MSA的技术提供了更准确和更有效的资源替代方案.
- 该研究鼓励社区专注于优化pLM基础模型,以实现更广泛的采用和可持续性.
相关概念视频
Conservation of Protein Domains Over Different Proteins
10.7K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.7K
From DNA to Protein
17.9K
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
17.9K
Conservation of Protein Domains
3.1K
3.1K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Protein Organization
6.2K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.2K
Protein Families
15.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.2K


