Related Experiment Video
Updated: Mar 3, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A method for learning a sparse classifier in the presence of missing data for high-dimensional biological datasets.
Kristen A Severson1, Brinda Monian1, J Christopher Love1
1Department of Chemical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA.
This study introduces expectation-maximization sparse discriminant analysis (EM-SDA) for building accurate and sparse classification models in biological and medical studies, effectively handling missing data. EM-SDA outperforms existing methods in accuracy and sparsity, even with incomplete datasets.
Area of Science:
- Biostatistics
- Bioinformatics
- Machine Learning
Background:
- Building classification models for biological and medical studies faces challenges with high-dimensional predictors and missing data.
- Supervised generative binary classification, specifically linear discriminant analysis (LDA), is a common approach.
- Existing methods struggle to simultaneously address sparsity and missing data effectively.
Purpose of the Study:
- To develop a novel algorithm, expectation-maximization sparse discriminant analysis (EM-SDA), for sparse LDA models.
- To effectively handle missing data within the classification modeling process.
- To improve the accuracy and sparsity of classification models in biomedical research.
Main Methods:
- Developed an expectation maximization (EM) algorithm to determine LDA parameters.
- Integrated priors into the EM algorithm to promote model sparsity.
- The EM-SDA algorithm handles both complete and incomplete datasets.
Main Results:
- EM-SDA demonstrated superior accuracy and sparsity compared to nearest shrunken centroids (NSCs) and sparse discriminant analysis (SDA) with imputation.
- Performance was evaluated through simulations and case studies on biomedical data.
- Results showed consistency between models trained on complete and incomplete data, highlighting robustness.
Conclusions:
- EM-SDA is an effective method for building sparse linear discriminant analysis models.
- The algorithm successfully addresses the dual challenges of feature selection (sparsity) and missing data imputation.
- EM-SDA offers a robust and accurate approach for classification in biomedical data analysis.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Biostatistics: Overview
Discrete variables are...
Survival Tree
Building a Survival Tree
Constructing a...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Classification of Systems-II

