Related Experiment Video
Updated: Nov 11, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
EPS: automated feature selection in case-control studies using extreme pseudo-sampling
Ruhollah Shemirani1, Stephane Wenric2, Eimear Kenny2
1Information Sciences Institute, University of Southern California, Marina del Rey, CA 90292, USA.
Summary:
Finding informative predictive features in high-dimensional biological case-control datasets is challenging. The Extreme Pseudo-Sampling (EPS) algorithm offers a solution to the challenge of feature selection via a combination of deep learning and linear regression models. First, using a variational autoencoder, it generates complex latent representations for the samples. Second, it classifies the latent representations of cases and controls via logistic regression. Third, it generates new samples (pseudo-samples) around the extreme cases and controls in the regression model. Finally, it trains a new regression model over the upsampled space. The most significant variables in this regression are selected. We present an open-source implementation of the algorithm that is easy to set up, use and customize. Our package enhances the original algorithm by providing new features and customizability for data preparation, model training and classification functionalities. We believe the new features will enable the adoption of the algorithm for a diverse range of datasets.
Availability And Implementation:
The software package for Python is available online at https://github.com/roohy/eps.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
Related Concept Videos
Censoring Survival Data
Randomized Experiments
Simple randomization
Simple...
Comparing the Survival Analysis of Two or More Groups
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Survival Tree
Building a Survival Tree
Constructing a...
Convenience Sampling Method
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...

