Related Experiment Video
Updated: Jul 4, 2025

10:34
Probing RNA Structure with Dimethyl Sulfate Mutational Profiling with Sequencing In Vitro and in Cells
Published on: December 9, 2022
4.2K
Optimal entropic properties of SARS-CoV-2 RNA sequences.
Marco Formentin1, Roberto Chignola2, Marco Favretti1
1Department of Mathematics Tullio Levi-Civita, University of Padova, via Trieste 63 35131 Padova, Italy.
Royal Society Open Science
|February 1, 2024
Summary
Scientists analyzed over a million SARS-CoV-2 genome sequences using information theory. This study reveals insights into the virus
Area of Science:
- Genomics
- Virology
- Information Theory
Background:
- The COVID-19 pandemic spurred rapid global collection of SARS-CoV-2 genome sequences.
- A large dataset of approximately 10^6 sequences is now available for analysis.
- The availability of a reference viral genome enables statistical studies of mutations.
Purpose of the Study:
- To apply information theory concepts to analyze the statistical properties of SARS-CoV-2 mutations.
- To quantify the information content and randomness of the viral mutation mechanism.
- To explore novel insights into viral evolution and adaptation.
Main Methods:
- Computation of Shannon entropy for each viral genome sequence.
- Calculation of relative entropy and mutual information between reference and mutated sequences.
- Statistical analysis of information-theoretic measures to understand mutation patterns.
Main Results:
- Information theory metrics reveal patterns in SARS-CoV-2 mutations.
- The study identifies known features, such as the CT bias, in a new format.
- Novel entropic properties of the mutation process were discovered, aligning with theoretical bounds.
Conclusions:
- Information theory provides a powerful framework for studying viral evolution.
- The SARS-CoV-2 mutation mechanism exhibits optimal entropic properties.
- This approach offers new perspectives on understanding viral genome dynamics and adaptation.
Related Concept Videos
RNA Stability
33.5K
Intact DNA strands can be found in fossils, while scientists sometimes struggle to keep RNA intact under laboratory conditions. The structural variations between RNA and DNA underlie the differences in their stability and longevity. Because DNA is double-stranded, it is inherently more stable. The single-stranded structure of RNA is less stable but also more flexible and can form weak internal bonds. Additionally, most RNAs in the cell are relatively short, while DNA can be up to 250 million...
33.5K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Viral Mutations
32.3K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
32.3K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
Bacterial RNA Polymerase
29.5K
Unlike eukaryotes, bacteria use a single RNA Polymerase (RNAP) to transcribe all genes. The different subunits of bacterial RNAPhave distinct functions. The multisubunit structure of the bacterial RNAP helps the enzyme to maintain catalytic function, facilitate assembly, interact with DNA and RNA, and self-regulate its activity.
In most genes, the transcription site is a single base present upstream of the coding sequence. Though RNAP is a catalytically efficient enzyme, it does not recognize...
In most genes, the transcription site is a single base present upstream of the coding sequence. Though RNAP is a catalytically efficient enzyme, it does not recognize...
29.5K
Nucleic Acid Structure
6.1K
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
DNA Structure
DNA...
6.1K

