Related Experiment Video
Updated: Apr 23, 2026

Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods
Published on: June 6, 2025
Annotated chemical patent corpus: a gold standard for text mining
Saber A Akhondi1, Alexander G Klenner2, Christian Tyrchan3
1Department of Medical Informatics, Erasmus University Medical Centre, Rotterdam, The Netherlands.
This study created a large, manually annotated patent corpus for chemical and biological entity recognition. This gold standard dataset aids in validating text mining tools for drug discovery research.
Area of Science:
- Medicinal Chemistry
- Bioinformatics
- Intellectual Property Analysis
Background:
- Patent analysis is vital for early-stage drug discovery, providing insights into prior art, novelty, and biological targets.
- Manual extraction of chemical and biological information from patents is time-consuming and resource-intensive.
- Automated text mining methods require robust, manually curated datasets for performance validation.
Purpose of the Study:
- To develop a large-scale, gold standard, manually annotated corpus of chemical patents.
- To establish annotation guidelines for chemical and biological entities within patents.
- To facilitate the validation of text mining tools for patent analysis in medicinal chemistry.
Main Methods:
- Selected 200 full patents from major patent offices (WIPO, USPTO, EPO).
- Developed annotation guidelines for chemicals, diseases, targets, and modes of action.
- Utilized automated pre-annotation followed by manual annotation by four independent groups, with a subset harmonized for inter-annotator agreement.
Main Results:
- Generated a comprehensive patent corpus with over 400,000 annotations.
- A harmonized subset of 47 patents yielded 36,537 annotations with derived agreement scores.
- The corpus includes annotations for chemicals, diseases, targets, modes of action, and OCR errors.
Conclusions:
- The created gold standard corpus is a valuable resource for training and validating text mining algorithms.
- This dataset will accelerate the extraction of crucial chemical and biological information from patents.
- Public availability of the corpus (www.biosemantics.org) supports advancements in drug discovery and intellectual property intelligence.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
NMR Spectroscopy of Aromatic Compounds
Aromatic Compounds: Overview
In 1825, Faraday...
NMR Spectroscopy Of Amines
NMR and Mass Spectroscopy of Carboxylic Acids
While α protons of carboxylic acids absorb at 2–2.5 ppm, β protons absorb further upfield.
Carboxylic acids are easily identified by dissolving them in deuterium oxide, which results in a rapid exchange of the acidic protons with deuterium. This leads to the...
Mass Spectrometry of Amines
Chemical Bonds
Atoms participate in a chemical bond formation to acquire a completed valence-shell electron configuration similar to that of the noble gas nearest to it in atomic number. Ionic, covalent, and metallic bonds are some of the important types of chemical bonds. Bond energy and bond length determine the strength of a chemical bond.
Types of Chemical Bonds
An ionic bond is formed due to electrostatic attraction between cations and anions. Often, the ions are formed by the transfer of electrons...