Related Experiment Video
Updated: Oct 14, 2025

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Automatic Extraction and Decryption of Abbreviations from Domain-Specific Texts
Michil Egorov1, Anastasia Funkner1
1ITMO University, Saint Petersburg, Russia.
Abstract:
This paper explores the problems of extraction and decryption of abbreviations from domain-specific texts in Russian. The main focus are unstructured electronic medical records which pose specific preprocessing problems. The major challenge is that there is no uniform way to write medical histories. The aim of the paper is to generalize the way of decrypting abbreviations from any variant of text. A dataset of nearly three million medical records was collected. A classifier model was trained in order to extract and decrypt abbreviations. After testing the proposed method with 224,307 records, the model showed an F1 score of 93.7% on a valid dataset.
Related Concept Videos
Extraction: Advanced Methods
Nomenclature of Alkynes
Nomenclature of Carboxylic Acid Derivatives: Acid Halides, Esters, and Acid Anhydrides
The IUPAC and common names of acid halides are derived from the corresponding carboxylic acids, by changing “ic acid” to “yl halide.” For example, as shown below, the IUPAC name ethanoyl chloride is derived from ethanoic acid, and the common name, acetyl chloride, is obtained from acetic acid.
IUPAC Nomenclature of Carboxylic Acids
For acyclic saturated monocarboxylic acids, the longest hydrocarbon chain containing the –COOH carbon is identified as the parent chain. Then, the last -e of the parent hydrocarbon name is replaced with a suffix -oic acid.
Nomenclature of Aromatic Compounds with a Single Substituent
IUPAC Nomenclature of Aldehydes

