Related Experiment Video
Updated: Jul 31, 2025

08:07
Author Spotlight: Advancing Antiviral Strategies Through Novel Immunocapture and Mass Spectrometry Techniques
Published on: January 12, 2024
795
Evaluating the effect of SARS-CoV-2 spike mutations with a linear doubly robust learner
Xin Wang1, Mingda Hu1, Bo Liu1
1Beijing Institute of Biotechnology, State Key Laboratory of Pathogen and Biosecurity, Beijing, China.
Frontiers in Cellular and Infection Microbiology
|May 8, 2023
Summary
Identifying key mutations in the SARS-CoV-2 Spike protein is crucial for understanding viral fitness. This study uses causal inference to pinpoint mutations that enhance viral fitness and predict transmission capacity.
Area of Science:
- Virology
- Genomics
- Computational Biology
Background:
- Diverse SARS-CoV-2 variants emerge due to Spike protein mutations, prolonging the pandemic.
- Identifying key mutations driving viral fitness is essential for pandemic control.
Purpose of the Study:
- To develop a causal inference framework for identifying key SARS-CoV-2 Spike mutations.
- To evaluate the contribution of mutations to viral fitness and predict transmission capacity.
Main Methods:
- Applied causal inference methods to large-scale SARS-CoV-2 genomes.
- Estimated statistical contributions of mutations to viral fitness across lineages.
- Validated identified mutations computationally for functional effects (stability, binding, immune escape).
Main Results:
- Identified key fitness-enhancing mutations (e.g., D614G, T478K) and critical protein regions (RBD, NTD).
- Developed mutational effect scores to compute strain fitness and predict transmission capacity.
- Validated prediction model with BA.2.12.1, demonstrating accuracy.
Conclusions:
- This is the first study applying causal inference to SARS-CoV-2 mutational analysis on large genomes.
- Findings provide systematic insights into SARS-CoV-2 evolution and guide functional studies of key mutations.
- The framework enables reliable prediction of viral transmission capacity from genomic sequences.
Related Concept Videos
Viral Mutations
32.7K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
32.7K
Residuals and Least-Squares Property
7.5K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.5K
Point and Frameshift Mutations
42
Point mutations are genetic alterations involving the change of a single nucleotide base pair in DNA. Depending on how the alteration affects protein synthesis, they can lead to various consequences.Point mutations fall into the following types:Silent mutations occur when a nucleotide change does not alter the amino acid sequence due to the redundancy of the genetic code. For instance, changing ACC to ACA still encodes threonine, leaving the protein function unaffected. This occurs because...
42
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K
Mutations
38.3K
Mutations are changes in the sequence of DNA. These changes can occur spontaneously or they can be induced by exposure to environmental factors. Mutations can be characterized in a number of different ways: whether and how they alter the amino acid sequence of the protein, whether they occur over a small or large area of DNA, and whether they occur in somatic cells or germline cells.
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
38.3K
Multiple Regression
3.1K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.1K

