Related Experiment Video
Updated: Jul 11, 2025

In Vivo Functional Study of Disease-associated Rare Human Variants Using Drosophila
Published on: August 20, 2019
MMPatho: Leveraging Multilevel Consensus and Evolutionary Information for Enhanced Missense Mutation Pathogenic
Fang Ge1,2, Muhammad Arif3,4, Zihao Yan5
1School of Geographic and Biologic Information, Nanjing University of Posts and Telecommunications, 9 Wenyuanlu, Nanjing 210023, China.
Abstract:
Understanding the pathogenicity of missense mutation (MM) is essential for shed light on genetic diseases, gene functions, and individual variations. In this study, we propose a novel computational approach, called MMPatho, for enhancing missense mutation pathogenic prediction. First, we established a large-scale nonredundant MM benchmark data set based on the entire Ensembl database, complemented by a focused blind test set specifically for pathogenic GOF/LOF MM. Based on this data set, for each mutation, we utilized Ensembl VEP v104 and dbNSFP v4.1a to extract variant-level, amino acid-level, individuals' outputs, and genome-level features. Additionally, protein sequences were generated using ENSP identifiers with the Ensembl API, and then encoded. The mutant sites' ESM-1b and ProtTrans-T5 embeddings were subsequently extracted. Then, our model group (MMPatho) was developed by leveraging upon these efforts, which comprised ConsMM and EvoIndMM. To be specific, ConsMM employs individuals' outputs and XGBoost with SHAP explanation analysis, while EvoIndMM investigates the potential enhancement of predictive capability by incorporating evolutionary information from ESM-1b and ProtT5-XL-U50, large protein language embeddings. Through rigorous comparative experiments, both ConsMM and EvoIndMM were capable of achieving remarkable AUROC (0.9836 and 0.9854) and AUPR (0.9852 and 0.9902) values on the blind test set devoid of overlapping variations and proteins from the training data, thus highlighting the superiority of our computational approach in the prediction of MM pathogenicity. Our Web server, available at http://csbio.njust.edu.cn/bioinf/mmpatho/, allows researchers to predict the pathogenicity (alongside the reliability index score) of MMs using the ConsMM and EvoIndMM models and provides extensive annotations for user input. Additionally, the newly constructed benchmark data set and blind test set can be accessed via the data page of our web server.
Insights
We developed MMPatho, a computational tool to predict missense mutation (MM) pathogenicity. MMPatho accurately identifies disease-causing mutations using variant and protein language model features, aiding genetic disease research.
Area of Science:
- Genomics and Bioinformatics
- Computational Biology
- Molecular Genetics
Background:
- Missense mutations (MMs) are crucial for understanding genetic diseases and individual variations.
- Accurate prediction of MM pathogenicity is essential but challenging.
Purpose of the Study:
- To develop a novel computational approach, MMPatho, for enhanced missense mutation pathogenicity prediction.
- To create robust benchmark and blind test datasets for evaluating MM pathogenicity prediction models.
Main Methods:
- Established a large-scale, nonredundant MM benchmark dataset and a focused blind test set.
- Extracted variant-level, amino acid-level, and genome-level features using Ensembl VEP and dbNSFP.
- Utilized protein sequence encoding and extracted embeddings from ESM-1b and ProtTrans-T5 for mutant sites.
- Developed two models, ConsMM (XGBoost with SHAP) and EvoIndMM (incorporating protein language embeddings).
Main Results:
- MMPatho models (ConsMM and EvoIndMM) achieved high performance on a blind test set (AUROC 0.9836-0.9854, AUPR 0.9852-0.9902).
- The models demonstrated superiority in predicting missense mutation pathogenicity.
- A web server was developed for public access to MMPatho prediction and data.
Conclusions:
- MMPatho offers a superior computational approach for predicting missense mutation pathogenicity.
- The developed datasets and web server facilitate further research in genetic disease and variant interpretation.
- Integrating evolutionary information and protein language embeddings significantly enhances predictive capabilities.
Related Concept Videos
Point and Frameshift Mutations
Mismatch Repair
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Evolutionary Relationships through Genome Comparisons
Spontaneous and Induced Mutations
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Mutations in Microorganisms

