Reliable prediction of Drosha processing sites improves microRNA gene prediction

Snorre A Helvik1, Ola Snøve, Pål Saetrom

  • 1Department of Computer and Information Science, Norwegian University of Science and Technology, NO-7052 Trondheim, Norway.

Abstract

Insights

We developed a new tool, Microprocessor SVM, to accurately identify microRNA (miRNA) processing sites. This aids in discovering new miRNA genes and verifying existing ones.

Area of Science:

  • Genomics
  • Molecular Biology
  • Bioinformatics

Background:

  • Mature microRNAs (miRNAs) are derived from hairpin precursors through enzymatic processing.
  • The initial Drosha processing step is crucial as it defines the mature miRNA product and is conserved across miRNA genes.
  • Accurate identification of Drosha processing sites is vital for discovering novel miRNA genes.

Purpose of the Study:

  • To develop a computational method for predicting Drosha processing sites in candidate miRNA hairpins.
  • To improve the accuracy and efficiency of miRNA gene discovery.
  • To re-evaluate existing miRNA annotations based on processing site prediction.

Main Methods:

  • Development of a Support Vector Machine (SVM) classifier, termed Microprocessor SVM, to predict 5' Drosha processing sites.
  • Training a secondary classifier on the output of Microprocessor SVM to enhance prediction of unconserved miRNAs.
  • Reanalysis of miRNA characteristics and supporting evidence for newly annotated miRNAs.

Main Results:

  • Microprocessor SVM correctly predicts the 5' Drosha processing site for 50% of known human 5' miRNAs.
  • 90% of Microprocessor SVM predictions are within two nucleotides of the true processing site.
  • A secondary classifier trained on Microprocessor SVM output outperforms existing methods for predicting unconserved miRNAs.

Conclusions:

  • The developed classifiers provide a robust method for identifying miRNA processing sites and discovering new miRNA genes.
  • Some previously annotated miRNAs may be misannotated, highlighting the need for experimental validation as Drosha and Dicer substrates.
  • The computational tools are publicly available for use in miRNA research.

Related Concept Videos

MicroRNAs01:22

MicroRNAs

MicroRNA (miRNA) are short, regulatory RNA transcribed from introns (non-coding regions of a gene) or intergenic regions (stretches of DNA present between genes). Several processing steps are required to form biologically active, mature miRNA. The initial transcript, called primary miRNA (pri-mRNA), base-pairs with itself, forming a stem-loop structure. Within the nucleus, an endonuclease enzyme, called Drosha, shortens the stem-loop structure into hairpin-shaped pre-miRNA. After the pre-miRNA...
MicroRNAs01:22

MicroRNAs

MicroRNA (miRNA) are short, regulatory RNA transcribed from introns—non-coding regions of a gene—or intergenic regions—stretches of DNA present between genes. Several processing steps are required to form biologically active, mature miRNA. The initial transcript, called primary miRNA (pri-mRNA), base-pairs with itself forming a stem-loop structure. Within the nucleus, an endonuclease enzyme, called Drosha, shortens the stem-loop structure into hairpin-shaped pre-miRNA. After the pre-miRNA ends...
MicroRNAs01:22

MicroRNAs

MicroRNA (miRNA) are short, regulatory RNA transcribed from introns—non-coding regions of a gene—or intergenic regions—stretches of DNA present between genes. Several processing steps are required to form biologically active, mature miRNA. The initial transcript, called primary miRNA (pri-mRNA), base-pairs with itself forming a stem-loop structure. Within the nucleus, an endonuclease enzyme, called Drosha, shortens the stem-loop structure into hairpin-shaped pre-miRNA. After the pre-miRNA ends...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...