通过贝叶斯主动学习和生物物理学预测高适应性病毒蛋白质变体
Marian Huot1,2, Dianzhuo Wang1,3, Jiacheng Liu4
1Department of Chemistry and Chemical Biology, Harvard University, Cambridge, MA 02138.
概括
早期检测高适应性病毒变种至关重要. 一个新的积极学习框架,VIRAL,加快了识别的五倍,使得迅速的流行病反应.
科学领域:
- 病毒学 病毒学
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 早期检测高适应性病毒变体对于疫情准备至关重要.
- 有限的实验资源往往阻碍了新兴变异的及时识别.
- 预测病毒健康和进化轨迹对于有效的公共卫生干预至关重要.
研究的目的:
- 引入VIRAL (通过快速主动学习进行病毒识别),这是一个主动学习框架,用于快速识别高适应性病毒变异.
- 通过计算预测和有限的实验验证,加快有关病毒变异的发现.
- 开发一个早期预警系统,对潜在的流行病病毒进行预警.
主要方法:
- 蛋白质语言模型,高斯过程与不确定性估计以及生物物理模型的整合.
- 应用少数射击学习来预测新型变体适应性.
- 与历史SARS-CoV-2数据和随机抽样策略进行基准测试.
主要成果:
- 与随机抽样相比,VIRAL可以加快高适应性变异的识别速度高达五倍.
- 该框架要求对不到1%的可能变体进行实验性表征.
- 识别了经常发生突变的部位,预测了抗体逃逸和ACE2结合的保存,提前长达两年.
- 不确定性驱动的变体选择促进了进化上遥远,潜在的危险变体的发现.
结论:
- VIRAL提供了一种计算高效和有效的方法,用于早期检测高适应性病毒变体.
- 该框架可以作为SARS-CoV-2和其他具有流行病潜力的新兴病毒的早期预警系统.
- 将预测建模与主动学习相结合,大大提高了变种监测的速度和范围.
相关概念视频
Leaky Scanning
5.2K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.2K
Physiological Pharmacokinetic Models: Assumption with Protein Binding
94
Physiological models with protein binding in pharmacokinetics offer a sophisticated approach to understanding drug disposition. These models consider drug-protein interactions, enabling them to effectively predict drug concentrations in different organs and tissues. This precision aids in accurate drug dosing, providing a significant advantage over conventional models. A key process within these models is equilibration, which ensures that drug concentrations achieve a steady state within the...
94
Viral Mutations
33.2K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
33.2K
Induced-fit Model
82.4K
Most chemical reactions in cells require enzymes—biological catalysts that speed up the reaction without being consumed or permanently changed. They reduce the activation energy needed to convert the reactants into products. Enzymes are proteins, that usually work by binding to a substrate—a reactant molecule that they act upon.
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
82.4K
Conservation of Protein Domains Over Different Proteins
11.5K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.5K


