评估进化信息对增强蛋白质语言模型嵌入的作用
Kyra Erckert1,2, Burkhard Rost3,4,5
1TUM School of Computation, Information and Technology, Bioinformatics & Computational Biology - i12, Boltzmannstr. 3, 85748, Garching/Munich, Germany. kyra.erckert@tum.de.
Scientific reports
|September 5, 2024
概括
蛋白质语言模型 (pLM) 的嵌入正在超越从多个序列对齐 (MSAs) 来进行蛋白质预测的传统进化信息. 将MSA与较新的pLM集成提供了有限的好处,并且在某些情况下可能会降低性能.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 结构生物学 结构生物学
背景情况:
- 蛋白质语言模型 (pLMs) 越来越多地用于蛋白质预测.
- 来自pLM的嵌入正在取代来自多重序列对齐 (MSA) 的进化信息.
- 目前正在调查pLM成功的原因,无论是固有的进化信息捕获还是其他因素.
研究的目的:
- 为了确定蛋白质语言模型 (pLM) 嵌入是否固有地捕获进化信息.
- 评估将进化信息 (来自MSA) 明确纳入pLM嵌入的影响.
- 为了比较基于pLM的方法与基于MSA的方法对蛋白质预测任务的性能.
主要方法:
- 测试各种方法,将进化信息集成到pLM嵌入式中.
- 在各种蛋白质预测任务中评估性能.
- 比较基于pLM,基于MSA和组合方法的方法.
主要成果:
- 较旧的pLM (SeqVec,ProtBert) 在纳入MSA时显示出显著的改善.
- 最近的pLM ProtT5没有从MSA集成中受益.
- 基于pLM的方法在各个任务中通常优于基于MSA的方法.
- 结合pLMs和MSAs有时会降低性能,特别是在内在疾病预测方面.
结论:
- pLM嵌入对于蛋白质预测非常有效,经常超过传统的基于MSA的方法.
- 从MSA中明确整合进化信息提供了有限的额外好处,特别是对于像ProtT5.5这样的先进的pLM.
- 基于pLM的方法的有效性表明,它们捕获的相关生物信息超出了明确的进化概况.
相关概念视频
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Improving Translational Accuracy
9.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.4K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
Proteins: From Genes to Degradation
12.1K
Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick. Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA...
Transcription is the synthesis of RNA...
12.1K


