Related Experiment Video
Updated: Aug 7, 2026

09:20
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Quantitative assessment of dictionary-based protein named entity tagging
Hongfang Liu1, Zhang-Zhi Hu, Manabu Torii
1Department of Biostatistics, Bioinformatics, and Biomathematics, Georgetown University Medical Center, Washington, DC 20007, USA. hl224@georgetown.edu
Summary
Biological named entity tagging (BNET) for protein names is complex due to high ambiguity and synonymy. BioThesaurus, a comprehensive gene/protein name resource, demonstrates high coverage of biological literature.
Area of Science:
- Bioinformatics
- Computational Biology
- Natural Language Processing
Background:
- Biological literature mining relies on accurate identification and normalization of gene and protein names.
- Natural language processing (NLP) offers methods for managing and extracting information from scientific texts.
- Biological Named Entity Tagging (BNET) is crucial for linking textual mentions to biological database entries.
Purpose of the Study:
- To quantitatively assess the complexity of BNET for protein entities.
- To evaluate the BioThesaurus resource for gene/protein name normalization.
- To analyze ambiguity, synonymy, and coverage of protein names within a biological thesaurus.
Main Methods:
- Evaluated BioThesaurus using metrics of ambiguity, synonymy, and coverage.
- Assessed name complexity both before and after normalization.
- Utilized the BioCreAtive dataset, comprising MEDLINE abstracts with gene/protein mentions.
Main Results:
- BioThesaurus contains over 2.6 million names, covering 1.8 million UniProtKB entries.
- Average synonymy decreased from 3.53 to 2.86 after normalization.
- Coverage reached 94.0% for gene/protein names in the BioCreAtive dataset, with ambiguity remaining high (around 2.3).
Conclusions:
- Gene and protein names exhibit significant ambiguity and multiple naming conventions.
- The BioThesaurus resource effectively covers the majority of gene/protein names found in biological literature.
- Normalization improves the precision of BNET by reducing synonymy.
Related Concept Videos
Tagging and Fusion Proteins
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
