通过深度学习和回归,从小数据中预测SARS-CoV-2尖端蛋白的演变
Samuel King1,2,3, Xinyi E Chen1,4,5, Sarah W S Ng1,4,5
1International Genetically Engineered Machine (iGEM) Team, University of British Columbia, Vancouver, BC, Canada.
Frontiers in systems biology
|August 14, 2025
概括
预测病毒演变对于公共卫生至关重要. 一个新的模型,VPRE,使用深度学习和回归来预测SARS-CoV-2尖端蛋白变化,但数据有限,识别潜在的新型变异.
科学领域:
- 病毒学 病毒学
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- SARS-CoV-2 变种推动了全球疫情爆发,使流行病控制复杂化.
- 病毒演变的预测模型需要大量的数据,这是新兴病毒的局限性.
研究的目的:
- 开发一种模型,用稀疏的数据来预测病毒蛋白的进化.
- 通过将序列编码为连续的数值表示来解决在建模离散突变中的计算挑战.
主要方法:
- 开发了一个病毒蛋白进化预测模型 (VPRE),结合了变化自编码器 (VAE) 和高斯过程 (GP) 回归.
- 使用VAE将离散的氨基酸序列编码成连续数.
- 在连续数值数据上使用GP回归来建模进化轨迹.
主要成果:
- 使用104个序列,VPRE成功预测了高达5个月的SARS-CoV-2尖端蛋白质演变.
- 预测包括了新的变异,最常见的预测与已知主导序列的单一氨基酸差异.
- 在尖端受体结合域 (RBD) 中预测的新型变异显示出强烈的结合人类ACE2 in silico.
结论:
- 结合深度学习和回归,可以利用稀疏的数据集进行有效的病毒蛋白进化建模.
- VPRE模型在预测未来的病毒变异及其潜在影响方面具有实用性.
- 这种方法支持开发更有效的医疗干预措施来对抗快速演变的病毒.
更多相关视频
相关概念视频
Steps in Outbreak Investigation
204
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
204
Single Nucleotide Polymorphisms-SNPs
15.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.9K
Viral Mutations
32.9K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
32.9K
End Point Prediction: Gran Plot
586
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
586


