Related Experiment Video
Updated: Jun 27, 2026

A Strategy for Sensitive, Large Scale Quantitative Metabolomics
Published on: May 27, 2014
Feature Down-Selection to Improve Supervised Classification by Machine Learning on Mass Spectrometry Imaging Data
Braysen Miller1, Aleesa E Chua1, Madeline Isom1
1Department of Chemistry, University of Kansas, 1450 Jayhawk Blvd, Lawrence, KS 66045, USA.
Abstract:
The advancements made in the mass spectrometry imaging (MSI) field have allowed for the generation of very large-scale data sets. These data are often interrogated by machine learning (ML), although storing and handling data sets of this size can be difficult. To aid impacted researchers, we seek to evaluate feature reduction strategies that will minimize the amount of data stored while still maintaining the ability to correctly classify the data. Two different feature selection strategies are tested on six different data sets, leveraging XGBoost as the machine learning algorithm. The study provides evidence that selecting features based on the greatest average abundance across all samples is best suited to scale down the feature set at a more modest trimming level, while selecting features based on statistical analysis via a Student's t-test is better suited for a more aggressive trimming level. These trends were present regardless of training set size or cross-validation strategy. The results from this work provide insight into when these feature filtering steps can be used effectively and when another data reduction strategy, including not restricting the data set, should be considered.
Related Concept Videos
MALDI-TOF Mass Spectrometry
Tandem Mass Spectrometry
Mass Spectrometry: Complex Analysis
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
High-Resolution Mass Spectrometry (HRMS)
Mass Spectrometry: Overview
Mass Spectrum: Interpretation
