Related Experiment Video
Updated: Jan 10, 2026

A Strategy for Sensitive, Large Scale Quantitative Metabolomics
Published on: May 27, 2014
A stable feature selection method based on majority voting and SHAP for high-dimensional metabolomics data
Zixuan Liu1, Jianqiang Du2, Jigen Luo3
1School of Intelligent Medicine and Information Engineering, Jiangxi University of Chinese Medicine, Nanchang 330004, China.
Background And Objective:
Metabolomics technology facilitates the simultaneous measurement of tens of thousands of metabolites, providing a crucial tool for disease mechanism research and biomarker screening. However, due to the high-dimensional and small-sample nature of the data, feature selection emerges as a critical step in data analysis. Existing feature selection methods still exhibit shortcomings in terms of stability, making it difficult to ensure consistency of selected features under sample perturbations or repeated experiments.
Methods:
Therefore, this study proposes a feature selection framework based on majority voting and SHAP integration (MVFS-SHAP), aimed at enhancing the stability and predictive performance of selected features. Specifically, five-fold cross-validation and bootstrap sampling techniques are used to generate multiple sampled datasets. The same base feature selection method is then applied to each dataset to generate corresponding feature subsets. Subsequently, a majority voting strategy is used to integrate these subsets. Based on Ridge regression and Linear SHAP, feature importance scores are computed, and features are re-ranked according to their average SHAP values. The top-ranked features are selected to form the final representative feature subset. Finally, a predictive model is constructed using partial least squares regression, and stability is evaluated through an extended Kuncheva index.
Results:
Extensive experiments on four high-dimensional, small-sample datasets demonstrate that the proposed MVFS-SHAP framework consistently outperforms existing feature aggregation strategies in terms of both feature selection stability and predictive accuracy. MVFS-SHAP shows high stability on most datasets. Among them, the stability of the Exo and Endo datasets exceeds 0.90, and about 80 % of the results are higher than 0.80. Even on challenging datasets, the stability remains within the range of 0.50 to 0.75. Furthermore, delivers lower RMSE values across models such as Lasso, Random Forest, and XGBoost. These results confirm the method's robustness in handling noisy and complex data while ensuring reliable biomarker selection.
Conclusions:
The proposed MVFS-SHAP framework provides a stable and effective solution for feature selection in high-dimensional, small-sample metabolomics studies. Future work will aim to enhance the framework's adaptability through optimized aggregation strategies and hyperparameter tuning, while extending its application to real-world biomedical research, including traditional Chinese medicine and precision medicine.

