Related Experiment Video
Updated: Jun 13, 2026

Lipidomics and Transcriptomics in Neurological Diseases
Published on: March 18, 2022
The Language of Elution: Autoregressive Prediction of the Next Feature in Untargeted LC-HRMS Lipidomics
1Department of Pharmacotherapy and Outcomes Sciences, Virginia Commonwealth University School of Pharmacy, Richmond, VA 23298, USA.
Abstract:
Untargeted liquid chromatography-high-resolution mass spectrometry (LC-HRMS) routinely detects thousands of molecular features per sample, yet only 2-20% receive confident structural annotations. A root cause of this "dark metabolome" is that tandem mass spectrometry (MS/MS) acquisition remains reactive: instruments select precursor ions after they appear, with no foreknowledge of what will elute next. Here we reframe chromatographic elution as an autoregressive sequence prediction task. Because reversed-phase elution order is governed by hydrophobicity, successive features are not independent draws but elements of a physically constrained sequence-analogous to tokens in natural language. We discretize the mass-to-charge (m/z) axis into 110 bins and train long short-term memory (LSTM) and Transformer models to predict the next eluting m/z bin from five per-token, annotation-free input features: m/z bin, mass defect, retention-time gap, ionization polarity, and intensity rank. Trained on 15,242 consensus features from four clinical lipidomics cohorts (342 human plasma samples, SCIEX TripleTOF 6600+, Waters CSH C18), the LSTM achieves 98.4% top-1 accuracy (99.99% top-5; mean absolute error = 3.6 Da) and the Transformer achieves 98.0% top-1 accuracy (99.88% top-5; MAE = 4 Da). Ablation analysis reveals that autoregressive sequence context accounts for 55.5 percentage points of accuracy; no individual input feature contributes more than 0.2 pp, establishing that the sequential pattern-not molecular properties-drives prediction. Cross-platform validation on an independent Agilent 6530 dataset using the same chromatographic method yields comparable performance (retention-time correlation r = 0.999), whereas datasets with different column chemistry (5.1% top-1) or different polarity acquisition mode (2.6% top-1, same instrument and column) fail catastrophically, confirming that models are specific to both chromatographic method and acquisition mode. However, full fine-tuning on as few as two to five quality-control injections recovers held-out analytical accuracy from 2.6% to nearly 50% top-1 (and to 99.6% on held-out QC), showing that cross-condition deployment is achievable with minimal calibration. A QC warm-up experiment reveals a null result: priming hidden states with quality-control injections confers no benefit over cold-start inference, indicating that the model captures elution logic from the sequence alone. Applying dual mass-plus-retention-time filtering to model predictions yields 168 putative annotations for previously unannotated features, illustrating potential for dark-lipidome characterization. These results establish that chromatographic elution sequences are highly predictable and lay the groundwork for predictive MS/MS acquisition strategies that could substantially improve annotation coverage in untargeted metabolomics.
More Related Videos
11:00Untargeted Metabolomics from Biological Sources Using Ultraperformance Liquid Chromatography-High Resolution Mass Spectrometry (UPLC-HRMS)
Published on: May 20, 2013
07:34Large Scale Non-targeted Metabolomic Profiling of Serum by Ultra Performance Liquid Chromatography-Mass Spectrometry (UPLC-MS)
Published on: March 14, 2013
Related Concept Videos
High-Performance Liquid Chromatography: Elution Process
High-Resolution Mass Spectrometry (HRMS)