Algorithms for detecting significantly mutated pathways in cancer

Fabio Vandin1, Eli Upfal, Benjamin J Raphael

  • 1Department of Computer Science, Brown University, Providence, Rhode Island, USA. vandinfa@cs.brown.edu

Insights

This study introduces a new network-based method to identify cancer-driving mutations. It helps distinguish functional mutations from passenger mutations by analyzing gene interactions, improving cancer pathway discovery.

Area of Science:

  • Genomics
  • Bioinformatics
  • Cancer Research

Background:

  • Cancer development is driven by somatic mutations across numerous genes, leading to mutational heterogeneity.
  • Distinguishing functional cancer mutations from passenger mutations is challenging.
  • Current methods often assess enrichment of mutated genes in known pathways.

Purpose of the Study:

  • To develop a novel computational approach for identifying cancer-driving subnetworks within gene interaction networks.
  • To provide an alternative to pathway enrichment analysis for understanding cancer mutations.
  • To efficiently identify statistically significant mutated subnetworks in a de novo manner.

Main Methods:

  • Utilized a genome-scale gene interaction network and somatic mutation data.
  • Employed a diffusion process on the network to define gene influence neighborhoods.
  • Implemented a two-stage multiple hypothesis test to control the false discovery rate (FDR).

Main Results:

  • Successfully identified known cancer-relevant pathways in glioblastoma and lung adenocarcinoma.
  • Discovered additional pathways implicated in other cancers but not previously reported in these samples.
  • Demonstrated the effectiveness of the network-based approach in uncovering novel cancer-associated subnetworks.

Conclusions:

  • The proposed network diffusion and hypothesis testing framework is an effective strategy for identifying cancer-driving subnetworks.
  • This method aids in distinguishing functional mutations and discovering novel cancer pathways.
  • The approach is computationally efficient and scalable for large cancer genome datasets.