Related Experiment Video
Updated: Jul 10, 2026

Chromatographic Fingerprinting by Template Matching for Data Collected by Comprehensive Two-Dimensional Gas Chromatography
Published on: September 2, 2020
Lossless compression of chemical fingerprints using integer entropy codes improves storage and retrieval.
Pierre Baldi1, Ryan W Benz, Daniel S Hirschberg
1Institute for Genomics and Bioinformatics, School of Information and Computer Sciences, University of California-Irvine, Irvine, CA 92697-3435, USA. pfbaldi@ics.uci.edu
This study introduces lossless compression for chemical fingerprints, reducing storage from over 1024 bits to ~300 bits per molecule. This enables exact similarity computations and improves retrieval performance for drug discovery.
Area of Science:
- Chemoinformatics
- Computational Chemistry
- Data Compression
Background:
- Modern chemoinformatics utilizes large fingerprint vectors for small molecules, often compressed lossily.
- Existing lossy compression methods can lead to information loss and affect similarity calculations.
Purpose of the Study:
- To develop efficient, lossless compression algorithms for chemical fingerprints.
- To enable exact similarity computations from compressed molecular representations.
- To improve retrieval performance in chemoinformatics databases.
Main Methods:
- Combined statistical fingerprint models with integer entropy codes (Golomb, Elias).
- Developed new lossless compression algorithms: monotone value (MOV) coding and monotone length (MOL) coding.
- Applied MOL Elias Gamma code to encode run lengths of reordered fingerprint components.
Main Results:
- Achieved lossless compression of chemical fingerprints to ~300 bits per molecule, nearing the Shannon entropy limit.
- Demonstrated that uncompressed similarity (Tanimoto) can be computed exactly from compressed data.
- Showcased significant retrieval performance improvements on six benchmark datasets of druglike molecules.
Conclusions:
- Lossless compression of chemical fingerprints is feasible and offers substantial storage reduction.
- The developed algorithms (MOV, MOL) provide efficient compression with modest computational cost.
- Exact similarity calculations from compressed data enhance chemoinformatics database retrieval efficiency.
Related Concept Videos
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
Entropy and Solvation
¹³C NMR: ¹H–¹³C Decoupling
A broadband decoupling technique is used to simplify these complex, sometimes overlapping, signals. Broadband decoupling relies on a...
¹H NMR: Interpreting Distorted and Overlapping Signals
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are slanted or...
Entropy
Third Law of Thermodynamics

