A graph theoretic approach to utilizing protein structure to identify non-random somatic mutations

Gregory A Ryslik1, Yuwei Cheng, Kei-Hoi Cheung

  • 1Department of Biostatistics, Yale School of Public Health, New Haven, CT, USA. gregory.ryslik@yale.edu.

BMC Bioinformatics
|March 28, 2014
PubMed
Abstract

Insights

A new method, GraphPAC, identifies cancer-driving mutations by analyzing protein structure. This approach reveals novel mutation clusters in key oncogenes and other proteins, improving cancer driver mutation discovery.

Area of Science:

  • Genomics
  • Bioinformatics
  • Structural Biology

Background:

  • Cancer development is linked to accumulated somatic mutations.
  • Driver mutations are key to tumorigenesis and therapeutic targets.
  • Current methods for driver mutation identification include machine learning and mutational clustering.

Purpose of the Study:

  • To introduce a novel methodology for identifying mutational clusters.
  • To leverage protein tertiary structure for enhanced driver mutation detection.
  • To improve the power of identifying potential driver mutations.

Main Methods:

  • Developed GraphPAC (Graph Protein Amino acid Clustering), a graph theoretical approach.
  • Integrated protein spatial structure data from the Protein Data Bank (PDB).
  • Utilized mutational data from the Catalogue of Somatic Mutations in Cancer (COSMIC).

Main Results:

  • GraphPAC identifies novel mutation clusters in known oncogenes like EGFR and KRAS.
  • Discovered new clusters in proteins such as DPP4 and NRP1, previously unidentified by other methods.
  • Demonstrated ability to detect clusters in proteins with and without prior evidence of clustering.

Conclusions:

  • GraphPAC offers an alternative and extension to existing methodologies like iPAC.
  • The graph theoretic approach effectively incorporates protein tertiary structure.
  • GraphPAC enhances the identification of potential activating driver mutations.

Related Concept Videos

Protein Networks02:26

Protein Networks

An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.7K
Conserved Binding Sites01:49

Conserved Binding Sites

Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.1K
Protein Organization01:24

Protein Organization

Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
7.2K
Protein-protein Interfaces02:04

Protein-protein Interfaces

Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Genetic Screens02:46

Genetic Screens

Genetic screens are tools used to identify genes and mutations responsible for phenotypes of interest. Genetic screens help identify individuals or a group of people at risk of developing  genetic diseases and help them with early intervention, targeted therapy, and reproductive options.
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
4.6K