Related Experiment Video
Updated: Apr 15, 2026

Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods
Published on: June 6, 2025
CHEMDNER: The drugs and chemical names extraction challenge
Martin Krallinger1, Florian Leitner2, Obdulia Rabal3
1Structural Computational Biology Group, Structural Biology and BioComputing Programme, Spanish National Cancer Research Centre, Calle Melchor Fernndez Almagro, 3, Madrid, Spain.
Abstract:
Natural language processing (NLP) and text mining technologies for the chemical domain (ChemNLP or chemical text mining) are key to improve the access and integration of information from unstructured data such as patents or the scientific literature. Therefore, the BioCreative organizers posed the CHEMDNER (chemical compound and drug name recognition) community challenge, which promoted the development of novel, competitive and accessible chemical text mining systems. This task allowed a comparative assessment of the performance of various methodologies using a carefully prepared collection of manually labeled text prepared by specially trained chemists as Gold Standard data. We evaluated two important aspects: one covered the indexing of documents with chemicals (chemical document indexing - CDI task), and the other was concerned with finding the exact mentions of chemicals in text (chemical entity mention recognition - CEM task). 27 teams (23 academic and 4 commercial, a total of 87 researchers) returned results for the CHEMDNER tasks: 26 teams for CEM and 23 for the CDI task. Top scoring teams obtained an F-score of 87.39% for the CEM task and 88.20% for the CDI task, a very promising result when compared to the agreement between human annotators (91%). The strategies used to detect chemicals included machine learning methods (e.g. conditional random fields) using a variety of features, chemistry and drug lexica, and domain-specific rules. We expect that the tools and resources resulting from this effort will have an impact in future developments of chemical text mining applications and will form the basis to find related chemical information for the detected entities, such as toxicological or pharmacogenomic properties.
More Related Videos
07:40A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
Published on: May 27, 2021
09:04Identifying Per- and Polyfluorinated Chemical Species with a Combined Targeted and Non-Targeted-Screening High-Resolution Mass Spectrometry Workflow
Published on: April 18, 2019
Related Concept Videos
Drug Nomenclature
Extraction: Advanced Methods
Nomenclature of Carboxylic Acid Derivatives: Acid Halides, Esters, and Acid Anhydrides
The IUPAC and common names of acid halides are derived from the corresponding carboxylic acids, by changing “ic acid” to “yl halide.” For example, as shown below, the IUPAC name ethanoyl chloride is derived from ethanoic acid, and the common name, acetyl chloride, is obtained from acetic acid.
Drug Discovery: Overview
Naming Enantiomers
Nomenclature of Alkanes
The alkane nomenclature considers the length of the carbon chain, the number, and the location of the substituent to arrive at its systematic name. The IUPAC...