Related Experiment Video
Updated: Sep 6, 2025

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
Published on: October 27, 2023
Chemical identification and indexing in PubMed full-text articles using deep learning and heuristics
Tiago Almeida1, Rui Antunes1, João F Silva1
1Department of Electronics, Telecommunications and Informatics (DETI), Institute of Electronics and Informatics Engineering of Aveiro (IEETA), University of Aveiro, Aveiro, Portugal.
Researchers developed an automated system for identifying chemicals in full-text biomedical articles, improving drug development research. The system enhances chemical identification, normalization, and Medical Subject Headings (MeSH) indexing for better information retrieval.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Drug Discovery
Background:
- Accurate chemical identification in biomedical literature is crucial for drug development.
- Previous research primarily focused on PubMed abstracts, neglecting valuable information in full-text articles.
- Manual indexing of Medical Subject Headings (MeSH) terms is time-consuming and requires expert knowledge.
Purpose of the Study:
- To develop and improve an automated system for chemical identification and MeSH indexing in PubMed full-text articles.
- To enhance the accuracy of chemical mention detection, entity normalization, and MeSH code assignment.
- To create a robust pipeline for processing biomedical literature and facilitating information retrieval.
Main Methods:
- A three-stage pipeline was implemented: chemical mention detection, entity normalization, and indexing.
- Deep learning models, including PubMedBERT, were utilized for chemical identification.
- A combination of dictionary filtering and deep learning similarity search was employed for normalization.
- Rule-based methods were developed for assigning relevant MeSH codes.
Main Results:
- The system achieved high performance in chemical identification (0.8731) and normalization (0.8275).
- Post-challenge improvements significantly boosted the overall system performance.
- The system demonstrated strong capabilities in indexing, achieving a score of 0.4849.
Conclusions:
- The developed system effectively automates chemical identification and MeSH indexing in full-text biomedical articles.
- The pipeline offers a valuable tool for researchers, improving the discoverability of chemical information.
- The publicly available code enables reproducibility and further development in the field.
Related Concept Videos
MALDI-TOF Mass Spectrometry
Matrix-assisted laser desorption ionization (MALDI) is a commonly...
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
NMR Spectroscopy of Aromatic Compounds
Mass Spectrometry: Complex Analysis
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
Methods of Classification and Identification
Mass Analyzers: Common Types

