EnsembleDesign: messenger RNA design minimizing ensemble free energy via probabilistic lattice parsing

Ning Dai1, Tianshuo Zhou1, Wei Yu Tang1

  • 1School of Electrical Engineering and Computer Science, Oregon State University, Corvallis, OR 97330, United States.

Abstract

Insights

Optimizing messenger RNA (mRNA) sequences requires considering all possible structures, not just one. Our new algorithm, EnsembleDesign, minimizes ensemble free energy for more flexible and stable mRNA designs.

Area of Science:

  • Computational Biology
  • Bioinformatics
  • Molecular Biology

Background:

  • Messenger RNA (mRNA) sequence design is crucial for applications like vaccines.
  • Previous methods focused on minimizing minimum free energy (MFE), neglecting alternative mRNA conformations.
  • Optimal mRNA function requires flexibility among multiple stable structures during translation.

Purpose of the Study:

  • To develop a novel algorithm for optimizing mRNA sequences by minimizing ensemble free energy.
  • To address the computational complexity of optimizing the entire Boltzmann ensemble of mRNA structures.

Main Methods:

  • Introduced EnsembleDesign, a novel algorithm using continuous relaxation.
  • Extended existing lattice representation and dynamic programming to probabilistic approaches.
  • Optimized expected ensemble free energy over a distribution of candidate sequences.

Main Results:

  • EnsembleDesign outperforms LinearDesign in minimizing ensemble free energy, particularly for longer sequences.
  • Ensemble designs exhibit lower average unpaired probabilities, reducing degradation.
  • Generated mRNA sequences demonstrate increased flexibility with flatter Boltzmann ensembles.

Conclusions:

  • Minimizing ensemble free energy is a more effective objective for mRNA sequence design.
  • EnsembleDesign provides a robust method for generating optimized mRNA sequences with improved stability and flexibility.

Related Concept Videos

Nonsense-mediated mRNA Decay02:27

Nonsense-mediated mRNA Decay

The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
10.9K
Leaky Scanning02:28

Leaky Scanning

During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA.  Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.2K
Nucleic Acid Structure01:25

Nucleic Acid Structure

The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms  a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
7.1K
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Regulation of Expression Occurs at Multiple Steps02:24

Regulation of Expression Occurs at Multiple Steps

Gene expression can be regulated at almost every step from gene to protein. Transcription is the step that is most commonly regulated. This involves the binding of proteins to short regulatory sequences on the DNA. This association can either promote or inhibit the transcription of a gene associated with the respective sequence.
Transcription results in the generation of precursor (pre-mRNA) that consists of both exons and introns, which needs further processing before being translated to a...
23.5K
RNA Splicing01:32

RNA Splicing

Splicing is the process by which eukaryotic RNA is edited before its translation into protein. The RNA strand transcribed from eukaryotic DNA is called the primary transcript. The primary transcripts that become mRNAs are called precursor messenger RNAs (pre-mRNAs). Eukaryotic pre-mRNA contains alternating sequences of exons and introns. Exons are nucleotide sequences that code for proteins, whereas introns are the non-coding regions. In RNA splicing, introns are removed and exons are bonded...
57.1K