Related Experiment Video
Updated: May 1, 2026

A Protocol for Functional Assessment of Whole-Protein Saturation Mutagenesis Libraries Utilizing High-Throughput Sequencing
Published on: July 3, 2016
A graph theoretic approach to utilizing protein structure to identify non-random somatic mutations
Gregory A Ryslik1, Yuwei Cheng, Kei-Hoi Cheung
1Department of Biostatistics, Yale School of Public Health, New Haven, CT, USA. gregory.ryslik@yale.edu.
Background:
It is well known that the development of cancer is caused by the accumulation of somatic mutations within the genome. For oncogenes specifically, current research suggests that there is a small set of "driver" mutations that are primarily responsible for tumorigenesis. Further, due to recent pharmacological successes in treating these driver mutations and their resulting tumors, a variety of approaches have been developed to identify potential driver mutations using methods such as machine learning and mutational clustering. We propose a novel methodology that increases our power to identify mutational clusters by taking into account protein tertiary structure via a graph theoretical approach.
Results:
We have designed and implemented GraphPAC (Graph Protein Amino acid Clustering) to identify mutational clustering while considering protein spatial structure. Using GraphPAC, we are able to detect novel clusters in proteins that are known to exhibit mutation clustering as well as identify clusters in proteins without evidence of prior clustering based on current methods. Specifically, by utilizing the spatial information available in the Protein Data Bank (PDB) along with the mutational data in the Catalogue of Somatic Mutations in Cancer (COSMIC), GraphPAC identifies new mutational clusters in well known oncogenes such as EGFR and KRAS. Further, by utilizing graph theory to account for the tertiary structure, GraphPAC discovers clusters in DPP4, NRP1 and other proteins not identified by existing methods. The R package is available at: http://bioconductor.org/packages/release/bioc/html/GraphPAC.html.
Conclusion:
GraphPAC provides an alternative to iPAC and an extension to current methodology when identifying potential activating driver mutations by utilizing a graph theoretic approach when considering protein tertiary structure.
Insights
A new method, GraphPAC, identifies cancer-driving mutations by analyzing protein structure. This approach reveals novel mutation clusters in key oncogenes and other proteins, improving cancer driver mutation discovery.
Area of Science:
- Genomics
- Bioinformatics
- Structural Biology
Background:
- Cancer development is linked to accumulated somatic mutations.
- Driver mutations are key to tumorigenesis and therapeutic targets.
- Current methods for driver mutation identification include machine learning and mutational clustering.
Purpose of the Study:
- To introduce a novel methodology for identifying mutational clusters.
- To leverage protein tertiary structure for enhanced driver mutation detection.
- To improve the power of identifying potential driver mutations.
Main Methods:
- Developed GraphPAC (Graph Protein Amino acid Clustering), a graph theoretical approach.
- Integrated protein spatial structure data from the Protein Data Bank (PDB).
- Utilized mutational data from the Catalogue of Somatic Mutations in Cancer (COSMIC).
Main Results:
- GraphPAC identifies novel mutation clusters in known oncogenes like EGFR and KRAS.
- Discovered new clusters in proteins such as DPP4 and NRP1, previously unidentified by other methods.
- Demonstrated ability to detect clusters in proteins with and without prior evidence of clustering.
Conclusions:
- GraphPAC offers an alternative and extension to existing methodologies like iPAC.
- The graph theoretic approach effectively incorporates protein tertiary structure.
- GraphPAC enhances the identification of potential activating driver mutations.
More Related Videos
08:46Implementation of In Vitro Drug Resistance Assays: Maximizing the Potential for Uncovering Clinically Relevant Resistance Mechanisms
Published on: December 9, 2015
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein-protein Interfaces
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...