Related Experiment Video
Updated: Sep 14, 2025

In Vivo Functional Study of Disease-associated Rare Human Variants Using Drosophila
Published on: August 20, 2019
Utilizing protein structure graph embeddings to predict the pathogenicity of missense variants
Martin Danner1,2, Matthias Begemann1, Miriam Elbracht1
1Institute for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, Aachen 52074, North-Rhine-Westphalia, Germany.
This study introduces a machine learning approach using protein structure predictions to determine the pathogenicity of missense variants. This method enhances variant interpretation and improves upon existing pathogenicity prediction scores.
Area of Science:
- Genomics and Bioinformatics
- Computational Biology
- Protein Structure Prediction
Background:
- Genetic variants, particularly missense variants, can alter protein structure and function, posing challenges for pathogenicity prediction.
- Existing bioinformatic tools for variant interpretation often do not directly incorporate protein structural information.
- Accurate prediction of missense variant pathogenicity is crucial for understanding genetic diseases.
Purpose of the Study:
- To develop a machine learning workflow for predicting the pathogenicity of missense variants using protein structure.
- To evaluate the utility of graph embeddings derived from predicted protein structures for pathogenicity prediction.
- To enhance existing pathogenicity prediction scores, such as CADD, using structural information.
Main Methods:
- Utilized the ESMFold protein language model to predict protein structures for missense variants.
- Employed graph autoencoders to generate embeddings from the predicted protein structures.
- Trained a classifier model using these graph embeddings to predict variant pathogenicity and compared performance with existing methods.
Main Results:
- Demonstrated that graph embeddings from predicted protein structures are effective for pathogenicity prediction.
- Showed that incorporating these graph embeddings can enhance the performance of the CADD score.
- Explored the impact of different graph embedding abstraction levels and compared embeddings from various protein-folding models.
Conclusions:
- The developed machine learning workflow provides a novel approach to pathogenicity prediction by leveraging predicted protein structures.
- Graph embeddings offer a valuable feature for improving the accuracy of missense variant interpretation.
- This method has the potential to refine our understanding of genetic variant effects on protein function and human health.
More Related Videos
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein and Protein Structures
Structural Protein Function