KinMutRF: a random forest classifier of sequence variants in the human protein kinase superfamily

Tirso Pons1, Miguel Vazquez1, María Luisa Matey-Hernandez2

  • 1Structural Biology and BioComputing Programme, Spanish National Cancer Research Centre (CNIO), Melchor Fernández Almagro, 3, 28029, Madrid, Spain.

BMC Genomics
|July 1, 2016
PubMed
Abstract

Insights

KinMutRF, a novel random-forest method, accurately identifies pathogenic variants in human kinases. This tool aids in understanding how genetic changes in protein kinases contribute to diseases like cancer.

Area of Science:

  • Genomics
  • Bioinformatics
  • Molecular Biology

Background:

  • Aberrant signaling by protein kinases is linked to human diseases, including cancer.
  • Identifying specific sequence variants in protein kinases that cause disease at a molecular level remains a challenge.
  • Most genomic alterations are tolerated by cells, with only a few disrupting molecular function and leading to disease.

Purpose of the Study:

  • To develop and evaluate KinMutRF, a novel random-forest method for automatically identifying pathogenic variants in human kinases.
  • To assess the performance of KinMutRF against existing methods using independent datasets.
  • To provide predictions for unclassified protein kinase variants.

Main Methods:

  • Developed KinMutRF, a random-forest classifier utilizing 26 decision trees.
  • Incorporated gene-level (Kinbase, Gene Ontology), domain-level (PFAM), and residue-level features (amino acid properties, functional annotations).
  • Trained and cross-validated KinMutRF on 3689 human kinase variants from UniProt, excluding unclassified variants.

Main Results:

  • KinMutRF achieved satisfactory performance in identifying disease-associated variants (Acc: 0.88, Prec: 0.82, Rec: 0.75, F-score: 0.78, MCC: 0.68).
  • The method demonstrated strong performance on independent kinase-specific mutation datasets (Kin-Driver, Pon-BTK).
  • Predictions were generated for 848 previously unclassified protein kinase variants.

Conclusions:

  • KinMutRF effectively classifies kinase variation with high performance.
  • Kinase-specific features significantly contribute to prediction accuracy, outperforming general methods.
  • The study advocates for the development of protein family-specific classifiers.

Related Concept Videos

Protein Kinases and Phosphatases02:54

Protein Kinases and Phosphatases

4.6K
Protein Kinases and Phosphatases02:54

Protein Kinases and Phosphatases

Proteins undergo chemical modifications that trigger changes in the charge, structure, and conformation of the proteins. Phosphorylation, acetylation, glycosylation, nitrosylation, ubiquitination, lipidation, methylation, and proteolysis are various protein modifications that regulate protein activity. Such modifications are usually enzyme-driven.
Protein kinases
Many proteins in the cell are regulated by phosphorylation, the addition of a phosphate group. A family of enzymes called kinases...
15.5K
Protein Families02:47

Protein Families

Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
17.4K
Protein Families02:47

Protein Families

4.6K
Signal Sequences and Sorting Receptors01:41

Signal Sequences and Sorting Receptors

Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
15.7K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
19.6K