Full-text chemical identification with improved generalizability and tagging consistency
Hyunjae Kim1, Mujeen Sung1, Wonjin Yoon1
1Department of Computer Science and Engineering, Korea University, Seoul, South Korea.
This study enhances chemical identification in full-text articles by improving named entity recognition and normalization. The developed model achieved top performance in a challenge, demonstrating superior accuracy for chemical entity tagging.
Area of Science:
- Biomedical Informatics
- Computational Chemistry
- Natural Language Processing
Background:
- Chemical identification in scientific literature is crucial for data mining and knowledge discovery.
- Existing models for chemical named entity recognition and normalization are primarily evaluated on titles and abstracts, with limited validation on full-text articles.
- Full-text analysis reveals limitations in current models, including poor generalizability to novel mentions and inconsistent tagging.
Purpose of the Study:
- To address the limitations of existing chemical identification models in full-text scientific articles.
- To develop and evaluate improved methods for chemical named entity recognition and normalization.
- To enhance the accuracy and reliability of extracting chemical information from full-text scientific documents.
Main Methods:
- Implemented transfer learning and mention-wise majority voting to improve generalizability and consistency in named entity recognition.
- Developed a hybrid model combining neural and dictionary-based approaches for named entity normalization, balancing recall and precision.
- Utilized simple training and post-processing techniques to refine model performance.
Main Results:
- Achieved first place in the BioCreative VII NLM-Chem challenge for named entity recognition with an 86.72 F1 score.
- Obtained an 78.31 F1 score for named entity normalization in the challenge, outperforming the median.
- Post-challenge re-implementation yielded an 84.70 F1 score for normalization, surpassing the challenge's best score by 3.34 F1 points.
Conclusions:
- The proposed methods effectively address limitations in chemical identification within full-text articles.
- The hybrid model demonstrates superior performance in both named entity recognition and normalization tasks.
- This work advances the state-of-the-art in automated chemical information extraction from scientific literature.
More Related Videos
05:35An Integrated Workflow of Identification and Quantification on FDR Control-Based Untargeted Metabolome
Published on: September 20, 2022
10:14Chromatographic Fingerprinting by Template Matching for Data Collected by Comprehensive Two-Dimensional Gas Chromatography
Published on: September 2, 2020
Related Concept Videos
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Chemical Shift: Internal References and Solvent Effects
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...
Tagging and Fusion Proteins
MALDI-TOF Mass Spectrometry
Matrix-assisted laser desorption ionization (MALDI) is a commonly...
Modern Molecular Taxonomy
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...
