Related Experiment Video
Updated: Aug 28, 2026

Single-throughput Complementary High-resolution Analytical Techniques for Characterizing Complex Natural Organic Matter Mixtures
Published on: January 7, 2019
From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for
Dipendra Bhandari1, Henry A Paz2, Keith Henderson1
1Arkansas Children's Nutrition Center and Arkansas Children's Research Institute, 15 Children's Way, Little Rock, AR, 72202, USA.
Introduction:
Untargeted metabolomics often results in a significant portion of unannotated metabolites, or "metabolic dark matter," which hinders biological interpretation.
Objectives:
A two-step analytical approach was developed to systematically prioritize and interpret unannotated metabolites using plasma LC-MS/MS data from pregnant women with obesity as a biologically relevant test dataset.
Methods:
The first step involved clustering 1,021 known metabolites into ten structurally coherent groups based on the Tanimoto similarity, thus defining the biologically relevant chemical space of the dataset. These metabolites were further characterized by Absorption, Distribution, Metabolism, and Excretion (ADME) profiling, protein target prediction, molecular docking and Kyoto Encyclopedia of Genes and Genomes pathway mapping analysis, to establish biological plausibility and functional perspective. Candidate structures for 1,836 unannotated features were retrieved from PubChem using molecular formula and molecular weight matching within a ±0.5 Da tolerance.
Results:
This search yielded 569,115 candidate structures, of which 368,197 unique structures were retained after curation. Tanimoto coefficient filtering reduced the candidate pool to 19,868 structurally plausible candidates, and retention time-based prioritization further refined this set to 418 high confidence candidate annotations, including 83 database-supported candidates identified through HMDB and LIPID MAPS structure database cross-referencing. RT-based prioritization effectively distinguished positional isomers sharing the same molecular formula by incorporating agreement between predicted and experimentally observed retention times.
Conclusion:
This improved discrimination among structurally similar candidates, expanded metabolite annotation confidence, and provided a scalable framework for prioritizing dark matter metabolites in untargeted metabolomics.

