Related Experiment Video
Updated: Jul 1, 2026

10:14
Chromatographic Fingerprinting by Template Matching for Data Collected by Comprehensive Two-Dimensional Gas Chromatography
Published on: September 2, 2020
Count your bits: fingerprint benchmarking to assess broad chemical space representation
Florian Huber1, Julian Pollmann2
1Centre for Digitalisation and Digitality, Düsseldorf University of Applied Sciences, Düsseldorf, Germany. florian.huber@hs-duesseldorf.de.
Journal of Cheminformatics
|June 30, 2026
Summary
Comparing molecular fingerprints reveals count and unfolded variants improve specificity and structural alignment. Folding can distort similarities, highlighting the need for unfolded representations and reproducible benchmarking tools like chemap.
Area of Science:
- Cheminformatics and computational chemistry.
- Development of molecular similarity metrics.
- Software development for cheminformatics research.
Background:
- Molecular similarity quantification is crucial for drug discovery and cheminformatics applications.
- Existing Tanimoto similarity measures using 2D fingerprints have limitations dependent on fingerprint type, representation, and folding.
- A systematic comparison of various fingerprint types and their representations is needed.
Purpose of the Study:
- To systematically compare common molecular fingerprint types and their representation variants.
- To evaluate fingerprint performance across diverse datasets using multiple criteria beyond retrieval.
- To introduce a standardized, open-source tool for reproducible fingerprint benchmarking.
Main Methods:
- Benchmarking of dictionary-based, circular (Morgan/FCFP), path-based (RDKit), topological-distance-based (Atom Pair), hybrid distance-encoded (MAP4), torsion, LINGO, and Avalon fingerprints.
- Evaluation using specificity, score distributions, compound-size dependence, and top-k ranking agreement on large, heterogeneous datasets.
- Comparison of fingerprint similarities against a graph-based reference.
Main Results:
- Count and log-count fingerprint variants generally enhance specificity and structural alignment compared to binary variants.
- Folding-induced bit collisions can significantly distort similarity scores, especially for high-occupancy fingerprints.
- Unfolded variants are crucial for certain fingerprints (e.g., RDKit) and often necessary for others (e.g., MAP4) on heterogeneous data.
Conclusions:
- Common default fingerprint settings, particularly folding, can introduce artifacts that negatively impact similarity assessments.
- Count and unfolded fingerprint representations offer improved specificity and better agreement with structure-based references.
- The open-source chemap library facilitates reproducible benchmarking and development of molecular fingerprints.
Related Concept Videos
IR Frequency Region: Fingerprint Region
IR spectra are divided into two main regions: the diagnostic region and the fingerprint region. The diagnostic region of the spectrum lies above 1500 cm−1. The absorptions resulting from single-bond vibrations of the N–H, C–H, and O–H stretch at higher wavenumbers and appear on the left side of the spectrum. The stretching absorptions of the C≡C and C≡N occur between 2100–2300 cm−1. In contrast, those arising from stretching absorptions of the C=O, C=N, and C=C occur between 1600–1850 cm−1.
The...
The...
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
