Detection of IUPAC and IUPAC-like chemical names
Roman Klinger1, Corinna Kolárik, Juliane Fluck
1Fraunhofer Institute Algorithms and Scientific Computing (SCAI), Department of Bioinformatics, Schloss Birlinghoven, 53574 Sankt Augustin, Germany. roman.klinger@scai.fraunhofer.de
Motivation:
Chemical compounds like small signal molecules or other biological active chemical substances are an important entity class in life science publications and patents. Several representations and nomenclatures for chemicals like SMILES, InChI, IUPAC or trivial names exist. Only SMILES and InChI names allow a direct structure search, but in biomedical texts trivial names and Iupac like names are used more frequent. While trivial names can be found with a dictionary-based approach and in such a way mapped to their corresponding structures, it is not possible to enumerate all IUPAC names. In this work, we present a new machine learning approach based on conditional random fields (CRF) to find mentions of IUPAC and IUPAC-like names in scientific text as well as its evaluation and the conversion rate with available name-to-structure tools.
Results:
We present an IUPAC name recognizer with an F(1) measure of 85.6% on a MEDLINE corpus. The evaluation of different CRF orders and offset conjunction orders demonstrates the importance of these parameters. An evaluation of hand-selected patent sections containing large enumerations and terms with mixed nomenclature shows a good performance on these cases (F(1) measure 81.5%). Remaining recognition problems are to detect correct borders of the typically long terms, especially when occurring in parentheses or enumerations. We demonstrate the scalability of our implementation by providing results from a full MEDLINE run.
Availability:
We plan to publish the corpora, annotation guideline as well as the conditional random field model as a UIMA component.
Related Concept Videos
Nomenclature of Alkanes
The alkane nomenclature considers the length of the carbon chain, the number, and the location of the substituent to arrive at its systematic name. The IUPAC...
IUPAC Nomenclature of Aldehydes
Nomenclature of Carboxylic Acid Derivatives: Amides and Nitriles
The IUPAC and common names of amides are derived from the parent carboxylic acid, by replacing the suffix “oic acid” and “ic acid,” respectively, with “amide.” In the following example, the IUPAC name ethanamide is derived from ethanoic acid, and the common name, acetamide, is obtained from acetic acid.
IUPAC Nomenclature of Carboxylic Acids
For acyclic saturated monocarboxylic acids, the longest hydrocarbon chain containing the –COOH carbon is identified as the parent chain. Then, the last -e of the parent hydrocarbon name is replaced with a suffix -oic acid.
Nomenclature of Carboxylic Acid Derivatives: Acid Halides, Esters, and Acid Anhydrides
The IUPAC and common names of acid halides are derived from the corresponding carboxylic acids, by changing “ic acid” to “yl halide.” For example, as shown below, the IUPAC name ethanoyl chloride is derived from ethanoic acid, and the common name, acetyl chloride, is obtained from acetic acid.
Common Names of Aldehydes and Ketones
Common names of aldehydes are derived from the names of their corresponding acid. For instance, the two-carbon aldehyde–acetaldehyde derives its name from the corresponding acid–acetic acid. Similarly, formaldehyde derives its name from formic acid and benzaldehyde from benzoic acid.
Aliphatic ketones are named by suffixing the word “ketone” to the alphabetically...


