Related Experiment Video
Updated: Jul 17, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Automatic extraction of protein point mutations using a graph bigram association.
Lawrence C Lee1, Florence Horn, Fred E Cohen
1Department of Cellular and Molecular Pharmacology, University of California San Francisco, San Francisco, California, United States of America.
Mutation GraB, a novel graph bigram algorithm, accurately extracts protein point mutations and their biological context from scientific literature. This method improves upon existing metrics for identifying mutation-protein associations, aiding research in protein structure and function.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Point mutations are crucial for understanding protein structure and function.
- Experimentally determined point mutations and their impacts are primarily documented in peer-reviewed literature.
- Existing databases often lack comprehensive indexing of point mutations and their associated biological context.
Purpose of the Study:
- To develop and evaluate an automated application, Mutation GraB (Graph Bigram), for identifying, extracting, and verifying point mutations from biomedical literature.
- To address the challenge of linking point mutations to their specific protein and organism of origin.
- To improve the accuracy of text-mining for mutation-protein associations.
Main Methods:
- Developed a graph-based bigram traversal algorithm to identify associations between point mutations, proteins, and organisms.
- Utilized the Swiss-Prot protein database for information verification.
- Incorporated term frequency and positional data within articles to enhance mutation-protein association.
- Compared the graph bigram metric against a word-proximity metric using full-text literature from GPCR, tyrosine kinase, and ion channel families.
Main Results:
- The graph bigram metric demonstrated superior performance over the word-proximity metric across all tested protein families (GPCRs, tyrosine kinases, ion channels).
- Achieved higher F-measures for GPCRs (0.79 vs. 0.76), protein tyrosine kinases (0.72 vs. 0.69), and ion channel transporters (0.76 vs. 0.74).
- Significantly improved precision in disambiguating point mutations associated with multiple proteins (0.84 vs. 0.73).
Conclusions:
- The graph bigram metric represents a significant advancement over previous search metrics for point mutation extraction.
- Mutation GraB offers a robust solution for text-mining applications requiring accurate association of words, particularly in the context of protein mutations.
- The developed method enhances the ability to mine biomedical literature for critical information on protein variations and their functional consequences.
More Related Videos
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
12:11Simultaneous Affinity Enrichment of Two Post-Translational Modifications for Quantification and Site Localization
Published on: February 27, 2020
Related Concept Videos
Point and Frameshift Mutations
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Mutations
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...