Related Experiment Video
Updated: Sep 26, 2025

Multi-step Preparation Technique to Recover Multiple Metabolite Compound Classes for In-depth and Informative Metabolomic Analysis
Published on: July 11, 2014
MetaboListem and TABoLiSTM: Two Deep Learning Algorithms for Metabolite Named Entity Recognition
Cheng S Yeung1, Tim Beck2,3, Joram M Posma1,3
1Section of Bioinformatics, Division of Systems Medicine, Department of Metabolism, Digestion and Reproduction, Faculty of Medicine, Imperial College London, London SW7 2AZ, UK.
Researchers developed two advanced methods for metabolite named entity recognition (NER) to improve metabolomics literature reviews. These deep learning models, MetaboListem and TABoLiSTM, achieve state-of-the-art performance in identifying metabolite names in scientific texts.
Area of Science:
- Bioinformatics
- Computational Biology
- Metabolomics
Background:
- The rapid growth of metabolomics literature presents challenges for manual literature reviews.
- Efficient text-mining tools are essential for navigating and analyzing this expanding body of research.
- Automated methods are needed to extract key information, such as metabolite names, from scientific publications.
Purpose of the Study:
- To develop and evaluate novel metabolite named entity recognition (NER) methods for the field of metabolomics.
- To create a standardized corpus of full-text metabolomics publications for training and evaluating NER models.
- To improve the efficiency and accuracy of literature reviews in metabolomics through advanced text-mining technologies.
Main Methods:
- Development of two metabolite NER models: MetaboListem (using GloVe embeddings) and TABoLiSTM (using BERT/BioBERT embeddings), both based on Bidirectional Long Short-Term Memory (BiLSTM) networks.
- Creation of a novel, standardized training corpus from over 1000 Open Access metabolomics publications, featuring 105,335 annotated metabolites.
- Evaluation of models on a manually annotated test corpus of 19,138 metabolite annotations.
Main Results:
- MetaboListem achieved an F1-score of 0.890 (precision 0.892, recall 0.888).
- TABoLiSTM (BioBERT version) demonstrated superior performance with an F1-score of 0.909 (precision 0.926, recall 0.893), achieving state-of-the-art results.
- The developed models accurately and efficiently identify metabolite names in text.
Conclusions:
- Deep learning algorithms, particularly BiLSTM networks with advanced transfer learning techniques, are highly effective for metabolite NER.
- The created corpus and NER algorithms provide valuable resources for various metabolomics text-mining applications, including information retrieval and literature-based discovery.
- These tools can significantly enhance the process of reviewing and analyzing metabolomics literature.
More Related Videos
07:34Large Scale Non-targeted Metabolomic Profiling of Serum by Ultra Performance Liquid Chromatography-Mass Spectrometry UPLC-MS
Published on: March 14, 2013
07:11Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
Published on: November 10, 2023
Related Concept Videos
Drug Metabolism: Phase II Reactions
Overview of Metabolism
Plant Metabolism
Sunlight, the primary source of energy in plants, is first absorbed by the chlorophyll pigments present in their leaves. Plants then use this energy to carry out photosynthesis, where water is oxidized into oxygen and carbon dioxide...
What is Metabolism?
Drug Metabolism: Phase I Reactions
Regulation of Metabolism