NTTMUNSW BioC modules for recognizing and normalizing species and gene/protein mentions
Hong-Jie Dai1, Onkar Singh2, Jitendra Jonnagaddala3
1Department of Computer Science and Information Engineering, National Taitung University, Taitung, Taiwan Interdisciplinary Program of Green and Information Technology, National Taitung University, Taitung, Taiwan hjdai@nttu.edu.tw emilysu@tmu.edu.tw.
This study introduces novel modules for normalizing species and gene names in biomedical literature, improving data accuracy for molecular interaction databases. These tools link gene variants to standardized NCBI Taxonomy IDs and Entrez Gene IDs.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Bioinformatics
Background:
- The rapid growth of biomedical literature presents challenges in accurately identifying and normalizing gene and species information.
- Ambiguity in gene and protein nomenclature complicates data curation for molecular interaction databases.
Purpose of the Study:
- To develop and evaluate a species normalization module that identifies species names and maps them to NCBI Taxonomy IDs.
- To create gene normalization modules for recognizing gene mentions and linking them to Entrez Gene IDs using a multistage algorithm.
- To ensure all developed modules are BioC-compatible and publicly available.
Main Methods:
- Developed a species normalization module that recognizes species names, including prefixes indicating gene origin, and normalizes them to NCBI Taxonomy IDs.
- Employed two separate modules for gene normalization, utilizing a multistage algorithm for processing full-text articles to identify gene mentions and map them to Entrez Gene IDs.
- Ensured modules are implemented as BioC-compatible .NET framework libraries.
Main Results:
- The species normalization module achieved a high F-score of 0.954 on an instance-level corpus.
- Successfully developed modules for both species and gene normalization, addressing complexities in biomedical text.
- All modules are publicly available on the NuGet gallery.
Conclusions:
- The developed modules effectively address the ambiguity in gene and species nomenclature within biomedical literature.
- These tools enhance the accuracy and efficiency of data curation for molecular interaction databases.
- The BioC-compatible libraries offer a valuable resource for the research community.
More Related Videos
Related Concept Videos
Modern Molecular Taxonomy
Genome Annotation and Assembly
Applications of Molecular Taxonomy
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Tagging and Fusion Proteins
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...


