Related Experiment Video
Updated: Jan 11, 2026

Biosynthesis of a Flavonol from a Flavanone by Establishing a One-pot Bienzymatic Cascade
Published on: August 14, 2019
Machine Learning for Group-Targeted Elution Order Prediction: Substituted Flavones as a Case Study
Ivan Rozanov1,2, Andrey Stavrianidi1,2, Aleksey Buryak1
1A.N. Frumkin Institute of Physical Chemistry and Electrochemistry, Russian Academy of Sciences, 31 Leninsky Prospect, GSP-1, Moscow 119071, Russia.
Abstract:
Prediction of the elution order of close structural analogs and isomers is a critical step in plant metabolites dereplication. The application of machine learning (ML) is an efficient approach to automate peak annotation by implementing structure-retention relationships. In this study, four ML models were trained to predict the elution order flavonoid derivative pairs containing a flavone aglycone bearing hydroxy and methoxy substituents. Additionally, retention times were indirectly estimated from model scores using linear regression and linear interpolation. Two data sets, comprising 51 compounds (1275 pairs) and 48 compounds (356 pairs), were constructed from available literature data. These data sets were used to explore alternative model training strategies and to conduct internal and external validation. A specially designed molecular fingerprint was employed to encode structural features of the flavone scaffold and its substituents, optimizing the representation for this widespread class of phytochemicals, taken as an example. Ranking neural network-based (NN) models with binary cross-entropy (BCE) and margin ranking (MR) loss functions employed a simplified version of the fingerprint, whereas logistic regression models were tested using both a condensed (20-bit) and an extended (92-bit) fingerprint that incorporated interaction between substituents. The pairwise error rates for elution order prediction were predominantly below 10%, demonstrating reliable performance under reversed-phase LC conditions with an acetonitrile gradient. Linear regression slightly outperformed the other models with the statistical significance indicated by Friedman and Wilcoxon tests. Although overall performance metrics were comparable, the use of a large uniform data set was found to be preferable over fragmented literature-derived data. The positive and negative effects of hydroxy and methoxy groups in different positions, along with their interactions on chromatographic retention, were analyzed following the visualization of model weights.
More Related Videos
09:04Identifying Per- and Polyfluorinated Chemical Species with a Combined Targeted and Non-Targeted-Screening High-Resolution Mass Spectrometry Workflow
Published on: April 18, 2019
08:56Detection of Regulated Ergot Alkaloids in Food Matrices by Liquid Chromatography-Trapped Ion Mobility Spectrometry-Time-of-Flight Mass Spectrometry
Published on: November 22, 2024
Related Concept Videos
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:
High-Performance Liquid Chromatography: Elution Process