Related Experiment Video
Updated: Jun 10, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Feature selection in finite mixture of sparse normal linear models in high-dimensional feature space
Abbas Khalili1, Jiahua Chen, Shili Lin
1Department of Mathematics and Statistics, McGill University, Montreal, Quebec H3A 2K6, Canada. khalili@math.mcgill.ca
This study introduces a new two-stage method for feature selection in complex genomics data. It effectively identifies key predictive variables within subpopulations, improving model accuracy and reducing false discoveries.
Area of Science:
- Genomics
- Statistical Modeling
- Bioinformatics
Background:
- Modern technology generates large, complex datasets, especially in genomics.
- Selecting relevant features from numerous variables is challenging with small sample sizes and multiple subpopulations.
- Accurate feature selection is crucial for reliable predictive modeling in high-dimensional biological data.
Purpose of the Study:
- To develop an effective feature selection method for finite mixture of sparse normal linear (FMSL) models.
- To address computational challenges and high false discovery rates in large feature spaces.
- To improve the accuracy of predictive models in genomics and related fields.
Main Methods:
- A two-stage procedure combining likelihood-based boosting and penalized likelihood methods.
- Likelihood-based boosting to reduce the dimensionality of candidate features.
- A novel scheme for initializing expectation-maximization estimation and an extended Bayesian information criterion for model selection.
Main Results:
- The proposed method successfully selects significant features while minimizing the inclusion of insignificant ones.
- Simulation studies demonstrate the procedure's effectiveness in feature selection.
- The method was applied to a real-world gene transcription regulation dataset.
Conclusions:
- The developed two-stage feature selection approach is effective for FMSL models in high-dimensional genomic data.
- The method overcomes computational hurdles and reduces false discovery rates.
- This approach enhances the reliability of predictive models in complex biological analyses.
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Expected Frequencies in Goodness-of-Fit Tests
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear.
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Frequency-dependent Selection
Outliers and Influential Points