Related Experiment Video
Updated: Jun 25, 2025

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
MoRF_ESM: Prediction of MoRFs in disordered proteins based on a deep transformer protein language model
Chun Fang1,2, Jiasheng He1, Hayato Yamana2
1Department of Information Engineering, Beijing Institute of Petrochemical Technology, 19 Qingyuan North Road, Daxing District, Beijing 102617, P. R. China.
Abstract:
Molecular recognition features (MoRFs) are particular functional segments of disordered proteins, which play crucial roles in regulating the phase transition of membrane-less organelles and frequently serve as central sites in cellular interaction networks. As the association between disordered proteins and severe diseases continues to be discovered, identifying MoRFs has gained growing significance. Due to the limited number of experimentally validated MoRFs, the performance of existing MoRF's prediction algorithms is not good enough and still needs to be improved. In this research, we present a model named MoRF_ESM, which utilizes deep-learning protein representations to predict MoRFs in disordered proteins. This approach employs a pretrained ESM-2 protein language model to generate embedding representations of residues in the form of attention map matrices. These representations are combined with a self-learned TextCNN model for feature extraction and prediction. In addition, an averaging step was incorporated at the end of the MoRF_ESM model to refine the output and generate final prediction results. In comparison to other impressive methods on benchmark datasets, the MoRF_ESM approach demonstrates state-of-the-art performance, achieving [Formula: see text] higher AUC than other methods when tested on TEST1 and achieving [Formula: see text] higher AUC than other methods when tested on TEST2. These results imply that the combination of ESM-2 and TextCNN can effectively extract deep evolutionary features related to protein structure and function, along with capturing shallow pattern features located in protein sequences, and is well qualified for the prediction task of MoRFs. Given that ESM-2 is a highly versatile protein language model, the methodology proposed in this study can be readily applied to other tasks involving the classification of protein sequences.
Insights
Identifying molecular recognition features (MoRFs) in disordered proteins is crucial for understanding disease. Our MoRF_ESM model, using deep learning protein representations, significantly improves MoRF prediction accuracy.
Area of Science:
- Biochemistry and Molecular Biology
- Computational Biology
- Protein Science
Background:
- Molecular recognition features (MoRFs) are key functional segments in intrinsically disordered proteins, vital for regulating membrane-less organelles and cellular interactions.
- The link between disordered proteins and diseases necessitates accurate MoRF identification, yet current prediction algorithms are limited by scarce experimental data.
Purpose of the Study:
- To develop an advanced computational model for predicting MoRFs in disordered proteins.
- To leverage deep learning protein representations for enhanced prediction accuracy.
Main Methods:
- Developed MoRF_ESM, a novel deep learning model integrating pretrained ESM-2 protein language model embeddings with a TextCNN architecture.
- Utilized attention map matrices from ESM-2 for residue representation and employed an averaging step for output refinement.
Main Results:
- MoRF_ESM achieved state-of-the-art performance on benchmark datasets, outperforming existing methods.
- Demonstrated significant improvements in AUC (Area Under the Curve) on TEST1 and TEST2 datasets compared to other approaches.
- Showcased the efficacy of combining deep evolutionary features from ESM-2 with shallow sequence patterns from TextCNN.
Conclusions:
- The MoRF_ESM model offers a powerful and accurate approach for MoRF prediction in disordered proteins.
- The methodology highlights the potential of integrating advanced protein language models with deep learning for various protein sequence classification tasks.
- This approach can be extended to other protein-related bioinformatics challenges, given the versatility of the ESM-2 model.
More Related Videos
Related Concept Videos
Intrinsically Disordered Proteins
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Conservation of Protein Domains
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Amyloid Fibrils
Amyloid deposits were observed as early as 1639 in the liver and the spleen. In 1854, Rudolph Virchow performed iodine staining,...

