Related Experiment Video
Updated: Sep 12, 2026

HPLC Coupled with Chemical Fingerprinting for Multi-Pattern Recognition for Identifying the Authenticity of Clematidis Armandii Caulis
Published on: November 11, 2022
Geographic authentication of premium tobacco extracts using machine learning models trained on LC-HRMS fingerprints
Xiu-Juan Xu1, Min Wang2, Xiao-Ping Li3
1Zhengzhou Tobacco Research Institute of China National Tobacco Corporation, Zhengzhou, China.
Introduction:
The geographic origin of premium tobacco extracts fundamentally shapes their sensory profiles and commercial value, but authenticating these complex matrices remains a formidable analytical challenge. Conventional nontargeted analysis workflows typically rely on a minute fraction of structurally annotated metabolites, thereby discarding the vast majority of unannotated yet highly specific chemical markers. To address this limitation, this study introduces a geographic authentication framework utilizing machine learning models trained on characteristic ion feature (CIF) fingerprints.
Methods:
A large-scale dataset of 1,697 tobacco samples was analyzed using liquid chromatography-high resolution mass spectrometry coupled with a data-independent acquisition strategy. Rigorous statistical filtering was applied to isolate CIFs and strip away the shared background metabolome. To evaluate the samples, univariate thresholding based on pattern similarity (dot product) and marker intensity ratios (R score) was initially tested. To overcome the vulnerability of univariate methods to batch effects and overlapping chemical profiles, optimized multivariate machine learning models were also deployed. The models were evaluated under a strict Leave-One-Batch-Out cross-validation framework to ensure robust cross-batch generalization.
Results:
Statistical filtering effectively stripped away over 90% of the shared background metabolome, isolating 47,543 CIFs for Zimbabwe tobacco and 3,646 CIFs for Yunnan tobacco. Univariate thresholding achieved a maximum balanced accuracy of 77.9% but proved too rigid. Under cross-validation, the multivariate Random Forest (RF) classifier demonstrated strong predictive performance, achieving a balanced accuracy of 74.6% for Zimbabwe tobacco and successfully detecting economically motivated adulteration across a wide dilution range (0-1,000 ppm). Conversely, both univariate and multivariate methods failed to achieve robust authentication for the Yunnan extracts. This underperformance was driven by severe intra-group heterogeneity stemming from diverse suppliers and heterogeneous processing treatments, which heavily diluted universal characteristic features.
Discussion:
Ultimately, this hybrid strategy provides a scalable tool for the industrial verification of consistent premium extracts. Furthermore, the challenges encountered with the Yunnan samples underscore the necessity of cohort-specific modeling when dealing with highly diverse agricultural products that are subjected to varied processing treatments.
