Related Experiment Video
Updated: Jun 30, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Distributionally Robust Feature Selection
Maitreyi Swaroop1, Tamar Krishnamurti2, Bryan Wilder1
1Machine Learning Department, Carnegie Mellon University.
None:
We study the problem of selecting limited features to observe such that models trained on them can perform well simultaneously across multiple subpopulations. This problem has applications in settings where collecting each feature is costly, e.g. requiring adding survey questions or physical sensors, and we must be able to use the selected features to create high-quality downstream models for different populations. Our method frames the problem as a continuous relaxation of traditional variable selection using a noising mechanism, without requiring back-propagation through model training processes. By optimizing over the variance of a Bayes-optimal predictor, we develop a model-agnostic framework that balances overall performance of downstream prediction across populations. We validate our approach through experiments on both synthetic datasets and real-world data.
Related Concept Videos
Frequency-dependent Selection
Quantifying and Rejecting Outliers: The Grubbs Test
Outliers and Influential Points
Expected Frequencies in Goodness-of-Fit Tests
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an organic...
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...