Related Experiment Video
Updated: Apr 15, 2026

Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods
Published on: June 6, 2025
Improving chemical entity recognition through h-index based semantic similarity.
Andre Lamurias1, João D Ferreira1, Francisco M Couto1
1LaSIGE, Departamento de Informática, Faculdade de Ciências, Universidade de Lisboa, 1749-016 Lisboa, Portugal.
This study introduces an improved semantic similarity measure for validating drug name recognition in text mining. The enhanced method boosts precision and recall, leading to better performance in the BioCreative IV CHEMDNER challenge.
Area of Science:
- Bioinformatics
- Natural Language Processing
- Computational Chemistry
Background:
- The CHEMDNER task in BioCreative IV focuses on recognizing and classifying drug names.
- Existing methods for chemical entity recognition often require robust validation techniques.
- Semantic similarity offers a promising avenue for validating these entities.
Purpose of the Study:
- To develop and evaluate an improved semantic similarity measure for validating chemical entities.
- To enhance the precision of drug name recognition in text mining.
- To adapt semantic similarity measures using the h-index for improved validation.
Main Methods:
- Applied semantic similarity validation techniques to Chemical Entities of Biological Interest (ChEBI) mappings.
- Adapted semantic similarity measures (simUI, simGIC) to incorporate the h-index of ancestors.
- Trained a Random Forest classifier incorporating semantic similarity scores.
- Compared adapted measures against their original versions for validation.
Main Results:
- The Random Forest classifier, using semantic similarity, improved F-measure by 4.6% over Conditional Random Fields classifiers.
- The h-index-based validation enhanced precision by reducing false positives at a fixed recall.
- Adapted measures based on the h-index demonstrated higher precision for equivalent recall levels.
Conclusions:
- The introduced semantic similarity measure is more efficient for validating text mining results.
- The enhanced validation improved recall and F-measure while maintaining high precision for the CHEMDNER task.
- Incorporating h-index into semantic similarity significantly boosts validation performance.
More Related Videos
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
09:04Identifying Per- and Polyfluorinated Chemical Species with a Combined Targeted and Non-Targeted-Screening High-Resolution Mass Spectrometry Workflow
Published on: April 18, 2019
Related Concept Videos
¹H NMR Chemical Shift Equivalence: Homotopic and Heterotopic Protons
Chemical Ionization (CI) Mass Spectrometry
¹H NMR Chemical Shift Equivalence: Enantiotopic and Diastereotopic Protons
In chiral compounds such as 2-butanol, replacing the methylene hydrogens at C3 produces a pair of...
Inductive Effects on Chemical Shift: Overview
Amines to Sulfonamides: The Hinsberg Test
Generally, a primary amine reacts with the Hinsberg reagent to produce an N-substituted benzenesulfonamide. The electron-withdrawing sulfonyl...
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...