Related Experiment Video
Updated: Jul 30, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
m6Aminer: Predicting the m6Am Sites on mRNA by Fusing Multiple Sequence-Derived Features into a CatBoost-Based
Ze Liu1, Pengfei Lan1, Ting Liu1,2
1College of Water Resources and Architectural Engineering, Northwest A&F University, Xianyang 712100, China.
Abstract:
As one of the most important post-transcriptional modifications, m6Am plays a fairly important role in conferring mRNA stability and in the progression of cancers. The accurate identification of the m6Am sites is critical for explaining its biological significance and developing its application in the medical field. However, conventional experimental approaches are time-consuming and expensive, making them unsuitable for the large-scale identification of the m6Am sites. To address this challenge, we exploit a CatBoost-based method, m6Aminer, to identify the m6Am sites on mRNA. For feature extraction, nine different feature-encoding schemes (pseudo electron-ion interaction potential, hash decimal conversion method, dinucleotide binary encoding, nucleotide chemical properties, pseudo k-tuple composition, dinucleotide numerical mapping, K monomeric units, series correlation pseudo trinucleotide composition, and K-spaced nucleotide pair frequency) were utilized to form the initial feature space. To obtain the optimized feature subset, the ExtraTreesClassifier algorithm was adopted to perform feature importance ranking, and the top 300 features were selected as the optimal feature subset. With different performance assessment methods, 10-fold cross-validation and independent test, m6Aminer achieved average AUC of 0.913 and 0.754, demonstrating a competitive performance with the state-of-the-art models m6AmPred (0.905 and 0.735) and DLm6Am (0.897 and 0.730). The prediction model developed in this study can be used to identify the m6Am sites in the whole transcriptome, laying a foundation for the functional research of m6Am.
Insights
This study introduces m6Aminer, a novel CatBoost-based tool for accurately identifying N6-methyladenosine (m6A) sites on mRNA. This method offers a faster, more efficient alternative to experimental techniques for understanding m6A
Area of Science:
- Bioinformatics
- Molecular Biology
- Computational Biology
Background:
- N6-methyladenosine (m6A) is a crucial post-transcriptional modification impacting mRNA stability and cancer progression.
- Accurate identification of m6A sites is vital for understanding its biological roles and medical applications.
- Current experimental methods for m6A site identification are costly and time-consuming, hindering large-scale analysis.
Purpose of the Study:
- To develop an efficient computational tool, m6Aminer, for identifying m6A sites on mRNA.
- To leverage machine learning, specifically CatBoost, for high-throughput m6A site prediction.
- To establish a foundation for comprehensive functional research of m6A modifications.
Main Methods:
- Employed nine distinct feature encoding schemes to construct an initial feature space for m6A site prediction.
- Utilized the ExtraTreesClassifier algorithm for feature importance ranking, selecting the top 300 features.
- Developed a CatBoost-based prediction model, m6Aminer, for identifying m6A sites.
Main Results:
- m6Aminer achieved an average AUC of 0.913 (10-fold cross-validation) and 0.754 (independent test).
- Demonstrated competitive performance compared to existing state-of-the-art models like m6AmPred and DLm6Am.
- The selected optimal feature subset significantly contributed to the model's predictive accuracy.
Conclusions:
- m6Aminer provides a robust and efficient computational approach for identifying m6A sites.
- The tool facilitates large-scale transcriptome-wide m6A site identification, accelerating research.
- This work lays the groundwork for deeper functional investigations into the role of m6A in biological processes and diseases.

