Related Experiment Video
Updated: May 24, 2026

14:18
A Strategy for Sensitive, Large Scale Quantitative Metabolomics
Published on: May 27, 2014
21.0K
Benchmarking feature selection and feature extraction methods to improve the performances of machine-learning
Justine Labory1,2,3, Evariste Njomgue-Fotso1, Silvia Bottini1,2
1Université Côte d'Azur, Center of Modeling Simulation and Interactions, Nice, France.
Computational and Structural Biotechnology Journal
|April 1, 2024
Summary
Feature selection enhances machine learning classification for omics data. Applying supervised feature selection improves performance in metabolomics, transcriptomics, and proteomics, aiding biomedical research.
Area of Science:
- Biomedical data analysis
- Machine learning in biology
Background:
- Omics data present challenges for machine learning due to heterogeneity, sparsity, and high dimensionality (p >> n).
- Class and feature imbalance are common issues in multi-omics datasets, hindering accurate classification.
- Metabolomics is crucial for understanding cancer, but its data complexity requires advanced analytical methods.
Purpose of the Study:
- To investigate the impact of feature extraction and selection techniques on machine learning classification performance with omics data.
- To develop guidelines for applying these techniques based on data characteristics.
- To evaluate the generalizability of these methods across different omics types.
Main Methods:
- Applied various linear and non-linear feature extraction techniques to three metabolomics datasets.
- Evaluated feature extraction methods with and without supervised feature selection.
- Tested the approach on transcriptomics and proteomics datasets to assess broader applicability.
Main Results:
- Supervised feature selection consistently improved the classification performance of feature extraction methods across all tested omics datasets.
- Demonstrated improved patient classification accuracy using the proposed feature engineering strategies.
- Provided a generalizable workflow and guidelines for optimizing omics data classification.
Conclusions:
- Feature selection is a critical step for enhancing machine learning model performance in omics data analysis.
- The findings offer practical strategies for improving classification accuracy in cancer research and other biomedical fields.
- The study provides valuable insights and tools for researchers working with complex biological datasets.

