Related Experiment Video
Updated: Jul 7, 2026

Simultaneous Affinity Enrichment of Two Post-Translational Modifications for Quantification and Site Localization
Published on: February 27, 2020
Cross-domain transfer learning from peptides to metabolites using a multi-property fine-tuned LLM
Uchenna Alex Anyaegbunam1, David Teschner2,3, Thierry Schmidlin4,5
1Computational Biology and Data Mining group (CBDM), Institute of Organismic and Molecular Evolution (iOME), Johannes Gutenberg University, Mainz, 55128, Germany.
Transfer learning using multi-task molecular property learning significantly improves retention time (RT) prediction for metabolites in metabolomics, especially with limited experimental data. This approach enhances compound identification accuracy in data-sparse conditions.
Area of Science:
- Computational Chemistry
- Cheminformatics
- Metabolomics
Background:
- Accurate retention time (RT) prediction is crucial for compound identification in metabolomics and lipidomics.
- Existing methods struggle with limited experimental RT data, hindering model generalization.
- Transfer learning offers a potential solution, but its application to metabolite RT prediction needs further investigation.
Purpose of the Study:
- To develop and evaluate a transfer learning framework for improving metabolite retention time (RT) prediction.
- To leverage large peptide datasets for enhanced RT prediction in data-sparse metabolomics conditions.
- To assess the impact of multi-task learning on the robustness and generalization of RT prediction models.
Main Methods:
- Developed a transfer learning framework using ChemBERTa, pretraining on large peptide datasets.
- Employed a multi-task learning objective to jointly predict RT and molecular descriptors.
- Evaluated model performance in low-data regimes for metabolite RT prediction.
Main Results:
- The multi-task pretrained model achieved superior generalization to metabolites compared to an RT-only model (median test R² of 0.842 vs. 0.820).
- Transfer learning significantly outperformed baseline models in low-data scenarios (e.g., 3% data: R² 0.322 vs. 0.216, MAE 114.9 vs. 131.7).
- Multi-task molecular property learning, not peptide pretraining alone, drove the performance gains.
Conclusions:
- Multi-task transfer learning is an effective and scalable strategy for enhancing metabolite RT prediction, particularly with limited experimental data.
- This approach improves compound identification accuracy in metabolomics and lipidomics.
- The developed framework provides a robust method for building comprehensive RT libraries.
Related Concept Videos
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
MALDI-TOF Mass Spectrometry
