从单个序列到进化轨迹:蛋白质语言模型捕捉了SARS-CoV-2的进化潜力
Kieran D Lamb1,2, Joseph Hughes1, Spyros Lytras1,3
1MRC-University of Glasgow Centre for Virus Research, School of Infection and Immunity, Glasgow, UK.
Nature communications
|February 19, 2026
概括
蛋白质语言模型 (PLMs) 分析蛋白质序列,以预测突变效应,而不需要多个序列对齐. 这些模型捕捉了进化史和变体特征,有助于理解病毒病原体.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 结构生物学是结构生物学.
背景情况:
- 蛋白质语言模型 (PLM) 利用自然语言处理概念来解释蛋白质序列.
- PLM可以从单独的氨基酸序列推断蛋白质结构和功能,绕过多重序列对齐 (MSA) 的需要.
研究的目的:
- 调查PLM表示,特别是ESM-2,如何评估突变诱导的蛋白质变异.
- 评估未经修改,预先训练的PLM在SARS-CoV-2的变异效应预测中的实用性.
主要方法:
- 在Silico深度突变扫描 (DMS) 用ESM-2 PLM对SARS-CoV-2尖端蛋白进行了.
- 评估了ESM-2直接从序列上下文中捕捉进化约束的能力.
- 该模型的性能与需要蛋白质结构或多个序列的方法进行了比较.
主要成果:
- ESM-2成功地从序列上下文捕获了进化约束,类似于基于MSA的方法.
- ESM-2表示编码了SARS-CoV-2变种的进化轨迹,并确定了令人担忧的变种的独特特征.
- 该模型展示了预测受体结合和抗原性转变的能力,并识别了表皮性相互作用.
结论:
- 像ESM-2这样的未经修改的预先训练的PLM是预测变异效应的强大工具,包括未观察到的突变.
- ESM-2提供了对病毒进化和病原体特征的洞察,适用于新型病毒病原体和任何蛋白质序列.
更多相关视频
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
相关概念视频
From DNA to Protein
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Proteins: From Genes to Degradation
Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick. Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA molecules by RNA...
Transcription is the synthesis of RNA molecules by RNA...
Leaky Scanning
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R stands for...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
