Related Experiment Video
Updated: Apr 23, 2026

Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods
Published on: June 6, 2025
A document classifier for medicinal chemistry publications trained on the ChEMBL corpus
George Papadatos1, Gerard Jp van Westen1, Samuel Croset1
1European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), European Molecular Biology Laboratory, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SD UK.
Background:
The large increase in the number of scientific publications has fuelled a need for semi- and fully automated text mining approaches in order to assist in the triage process, both for individual scientists and also for larger-scale data extraction and curation into public databases. Here, we introduce a document classifier, which is able to successfully distinguish between publications that are 'ChEMBL-like' (i.e. related to small molecule drug discovery and likely to contain quantitative bioactivity data) and those that are not. The unprecedented size of the medicinal chemistry literature collection, coupled with the advantage of manual curation and mapping to chemistry and biology make the ChEMBL corpus a unique resource for text mining.
Results:
The method has been implemented as a data protocol/workflow for both Pipeline Pilot (version 8.5) and KNIME (version 2.9) respectively. Both workflows and models are freely available at: ftp://ftp.ebi.ac.uk/pub/databases/chembl/text-mining. These can be readily modified to include additional keyword constraints to further focus searches.
Conclusions:
Large-scale machine learning document classification was shown to be very robust and flexible for this particular application, as illustrated in four distinct text-mining-based use cases. The models are readily available on two data workflow platforms, which we believe will allow the majority of the scientific community to apply them to their own data.
ᅟ:
Graphical AbstractMultidimensional scaling analysis applied to document vectors derived from titles and abstracts in different corpora. Notably, there is large overlap between the documents in the different ChEMBL versions and BindingDB, while the background MEDLINE set is largely divergent.
More Related Videos
07:29HPLC Coupled with Chemical Fingerprinting for Multi-Pattern Recognition for Identifying the Authenticity of Clematidis Armandii Caulis
Published on: November 11, 2022
14:34A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Related Concept Videos
Classification of Neurotransmitters
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...
Classification of Elements and Compounds
Compounds are pure substances composed of two or more elements in fixed, definite proportions. Compounds are classified as ionic or molecular (covalent) based on the bonds...
Molecules with Multiple Chiral Centers
MALDI-TOF Mass Spectrometry
Drug Discovery: Overview