Related Experiment Video
Updated: Feb 25, 2026

A Complete Pipeline for Isolating and Sequencing MicroRNAs, and Analyzing Them Using Open Source Tools
Published on: August 21, 2019
Improving the Quality of Positive Datasets for the Establishment of Machine Learning Models for pre-microRNA
Abstract:
MicroRNAs (miRNAs) are involved in the post-transcriptional regulation of protein abundance and thus have a great impact on the resulting phenotype. It is, therefore, no wonder that they have been implicated in many diseases ranging from virus infections to cancer. This impact on the phenotype leads to a great interest in establishing the miRNAs of an organism. Experimental methods are complicated which led to the development of computational methods for pre-miRNA detection. Such methods generally employ machine learning to establish models for the discrimination between miRNAs and other sequences. Positive training data for model establishment, for the most part, stems from miRBase, the miRNA registry. The quality of the entries in miRBase has been questioned, though. This unknown quality led to the development of filtering strategies in attempts to produce high quality positive datasets which can lead to a scarcity of positive data. To analyze the quality of filtered data we developed a machine learning model and found it is well able to establish data quality based on intrinsic measures. Additionally, we analyzed which features describing pre-miRNAs could discriminate between low and high quality data. Both models are applicable to data from miRBase and can be used for establishing high quality positive data. This will facilitate the development of better miRNA detection tools which will make the prediction of miRNAs in disease states more accurate. Finally, we applied both models to all miRBase data and provide the list of high quality hairpins.
Insights
This study developed machine learning models to assess the quality of microRNA (miRNA) data, ensuring more accurate disease prediction. The models identify high-quality miRNA hairpins, improving computational detection tools.
Area of Science:
- Bioinformatics
- Genomics
- Molecular Biology
Background:
- MicroRNAs (miRNAs) regulate protein abundance and influence phenotypes, implicating them in diseases like cancer.
- Computational methods using machine learning are crucial for pre-miRNA detection due to experimental complexity.
- The quality of miRNA data in miRBase, a key resource, is uncertain, necessitating filtering strategies that can reduce data availability.
Purpose of the Study:
- To develop and validate machine learning models for assessing the quality of pre-miRNA datasets.
- To identify features that distinguish between low and high-quality pre-miRNA data.
- To apply these models to miRBase data to curate a high-quality set of miRNA hairpins.
Main Methods:
- Developed a machine learning model to evaluate pre-miRNA data quality based on intrinsic measures.
- Analyzed features of pre-miRNAs to differentiate between low and high-quality data.
- Applied the developed models to all entries in miRBase.
Main Results:
- The machine learning model effectively established pre-miRNA data quality.
- Specific features were identified that discriminate between low and high-quality datasets.
- A curated list of high-quality hairpins from miRBase was generated using the developed models.
Conclusions:
- The developed models provide a reliable method for data quality assessment in miRNA research.
- These models facilitate the creation of high-quality positive datasets, essential for training accurate miRNA detection tools.
- The curated list of high-quality hairpins will advance miRNA prediction in disease states.
Related Concept Videos
MicroRNAs
MicroRNAs

