m6Aminer: Predicting the m6Am Sites on mRNA by Fusing Multiple Sequence-Derived Features into a CatBoost-Based

Ze Liu1, Pengfei Lan1, Ting Liu1,2

  • 1College of Water Resources and Architectural Engineering, Northwest A&F University, Xianyang 712100, China.

Insights

This study introduces m6Aminer, a novel CatBoost-based tool for accurately identifying N6-methyladenosine (m6A) sites on mRNA. This method offers a faster, more efficient alternative to experimental techniques for understanding m6A

Area of Science:

  • Bioinformatics
  • Molecular Biology
  • Computational Biology

Background:

  • N6-methyladenosine (m6A) is a crucial post-transcriptional modification impacting mRNA stability and cancer progression.
  • Accurate identification of m6A sites is vital for understanding its biological roles and medical applications.
  • Current experimental methods for m6A site identification are costly and time-consuming, hindering large-scale analysis.

Purpose of the Study:

  • To develop an efficient computational tool, m6Aminer, for identifying m6A sites on mRNA.
  • To leverage machine learning, specifically CatBoost, for high-throughput m6A site prediction.
  • To establish a foundation for comprehensive functional research of m6A modifications.

Main Methods:

  • Employed nine distinct feature encoding schemes to construct an initial feature space for m6A site prediction.
  • Utilized the ExtraTreesClassifier algorithm for feature importance ranking, selecting the top 300 features.
  • Developed a CatBoost-based prediction model, m6Aminer, for identifying m6A sites.

Main Results:

  • m6Aminer achieved an average AUC of 0.913 (10-fold cross-validation) and 0.754 (independent test).
  • Demonstrated competitive performance compared to existing state-of-the-art models like m6AmPred and DLm6Am.
  • The selected optimal feature subset significantly contributed to the model's predictive accuracy.

Conclusions:

  • m6Aminer provides a robust and efficient computational approach for identifying m6A sites.
  • The tool facilitates large-scale transcriptome-wide m6A site identification, accelerating research.
  • This work lays the groundwork for deeper functional investigations into the role of m6A in biological processes and diseases.