Related Experiment Video
Updated: Jun 23, 2026

14:18
A Strategy for Sensitive, Large Scale Quantitative Metabolomics
Published on: May 27, 2014
21.0K
METAbolomics data Balancing with Over-sampling Algorithms (META-BOA): an online resource for addressing class
Emily Hashimoto-Roth1,2, Anuradha Surendra3, Mathieu Lavallée-Adam1
1Department of Biochemistry, Microbiology and Immunology, Ottawa Institute of Systems Biology, Ottawa, ON, Canada.
Bioinformatics (Oxford, England)
|October 12, 2022
Summary
Class imbalance in metabolomic and lipidomic data mining is addressed by META-BOA, a web application simplifying over-sampling algorithm evaluation. It helps users balance datasets and visualize augmentation effects for improved machine learning model performance.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning
Background:
- Class imbalance is a significant challenge in metabolomic and lipidomic data mining, potentially leading to overfitting.
- Existing methods for handling class imbalance lack user-friendly accessibility for those with limited computational expertise.
- There is a need for a platform to easily evaluate the impact of various over-sampling techniques on data.
Purpose of the Study:
- To introduce META-BOA, a web-based application designed to address class imbalance in omics data.
- To provide users with an accessible tool for selecting, applying, and evaluating different over-sampling algorithms.
- To facilitate the comparison of data before and after over-sampling to understand augmentation effects.
Main Methods:
- META-BOA offers four over-sampling methods: Synthetic Minority Over-sampling Technique (SMOTE), Borderline-SMOTE (BSMOTE), Adaptive Synthetic (ADASYN), and Random Over-Sampling Examples (ROSE).
- The application generates a balanced dataset by creating synthetic samples for the minority class.
- Principal Component Analysis (PCA) and t-distributed stochastic neighbor embedding (t-SNE) visualizations are used to display data pre- and post-over-sampling.
Main Results:
- META-BOA successfully generates balanced datasets using selected over-sampling techniques.
- Visualizations (PCA and t-SNE) effectively demonstrate the impact of over-sampling on data distribution.
- Random forest classification comparison highlights differences in model performance on original versus balanced datasets.
Conclusions:
- META-BOA provides a valuable and accessible resource for researchers dealing with class imbalance in metabolomic and lipidomic data.
- The tool aids in selecting appropriate over-sampling strategies by allowing direct evaluation of their effects.
- Improved data balancing through META-BOA can lead to more robust and reliable machine learning models in omics research.
More Related Videos
Related Concept Videos
Genomics
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
Biostatistics: Overview
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...

