Related Experiment Video
Updated: Jul 7, 2026

Simultaneous Affinity Enrichment of Two Post-Translational Modifications for Quantification and Site Localization
Published on: February 27, 2020
Cross-domain transfer learning from peptides to metabolites using a multi-property fine-tuned LLM
Uchenna Alex Anyaegbunam1, David Teschner2,3, Thierry Schmidlin4,5
1Computational Biology and Data Mining group (CBDM), Institute of Organismic and Molecular Evolution (iOME), Johannes Gutenberg University, Mainz, 55128, Germany.
Motivation:
Accurate liquid chromatography retention time (RT) prediction is a critical component of compound identification in metabolomics and lipidomics. However, existing RT prediction approaches are often limited by the scarcity of experimental RT measurements for many molecular classes, restricting model generalization and the construction of comprehensive RT libraries. Transfer learning from data-rich chemical domains offers a potential strategy to overcome these limitations, but its effectiveness for metabolite RT prediction remains insufficiently explored.
Results:
We developed a transfer learning framework based on ChemBERTa that leverages large peptide datasets to improve metabolite RT prediction under data-sparse conditions. A peptide-pretrained model was trained using a multi-task objective that jointly predicted RT and seven RDKit-derived molecular descriptors. Compared with an RT-only model, the multi-task approach learned more robust chemical representations and demonstrated superior generalization to metabolites, achieving a median test R² of 0.842 versus 0.820. When transferred to metabolite RT prediction, the multi-task pretrained model substantially outperformed models trained from scratch at low-data regimes. Using only 3% of metabolite training data (2129 compounds), transfer learning achieved a median test R² of 0.322 compared with 0.216 for the baseline model, while reducing MAE from 131.7 to 114.9. Significant improvements were also observed at 5% and 10% training fractions, with benefits gradually diminishing as larger metabolite datasets became available. In contrast, a peptide-pretrained single-task RT model showed performance comparable to the baseline, indicating that the observed gains arise primarily from multi-task molecular property learning rather than peptide pretraining alone. These findings demonstrate that multi-task transfer learning provides an effective and scalable strategy for improving RT prediction in metabolomics, particularly when experimental training data are limited.
Availability:
Freely available on https://github.com/uchealex/CHEMBEDDING.
Related Concept Videos
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
MALDI-TOF Mass Spectrometry
