Related Experiment Video
Updated: Sep 11, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Multivariate Optimization of k for k-Nearest-Neighbor Feature Selection With Dichotomous Outcomes: Complex
New methods for nearest-neighbor feature selection improve detection of relevant features, especially for imbalanced data. These techniques enhance Nearest-neighbor Projected-Distance Regression and Relief-based algorithms, showing promise in complex datasets like Major Depressive Disorder RNASeq data.
Area of Science:
- Machine Learning
- Bioinformatics
- Computational Biology
Background:
- Nearest-neighbor feature selection is crucial but sensitive to data characteristics like sample size, feature count, statistical effects, and class imbalance.
- Existing methods, including Nearest-neighbor Projected-Distance Regression (NPDR), often do not adequately address class imbalance.
- Effective feature selection is vital for accurate biological data analysis and disease-related gene discovery.
Purpose of the Study:
- To develop and evaluate novel fixed-k strategies for nearest-neighbor feature selection that explicitly handle class imbalance.
- To improve the performance of NPDR and Relief-based algorithms on imbalanced datasets.
- To compare the efficacy of these new methods against established techniques like random forest and ridge regression.
Main Methods:
- Introduced minority-class-k for NPDR, parameterizing fixed-k by minority class size.
- Developed hit-miss-k, a class-adaptive fixed-k for Relief-based algorithms.
- Presented variable-wise optimized k (VWOK) and principal components analysis-derived k (kPCA) for imbalance-adaptive optimization.
- Utilized simulated data and consensus-nested cross-validation (cnCV) for rigorous performance evaluation.
Main Results:
- The proposed methods significantly enhanced feature detection across various nearest-neighbor scoring metrics.
- Novel fixed-k strategies demonstrated superior performance compared to random forest and ridge regression.
- Application to Major Depressive Disorder RNASeq data revealed NPDR with minority-class-k identified highly relevant genes for the condition.
Conclusions:
- The developed fixed-k methods effectively address class imbalance in nearest-neighbor feature selection.
- These adaptive strategies offer improved accuracy and relevance in feature discovery, particularly for imbalanced biological datasets.
- NPDR with minority-class-k shows potential as a reliable tool for identifying biologically significant genes in complex disease studies.
Related Concept Videos
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Depressive Disorders: MDD and Dysthymia
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Depressive Disorders: Etiology
Biological Factors in Depression
Biological predispositions significantly influence the risk of developing depressive disorders. Genetic studies highlight the role of variations in the serotonin transporter...
Cancer Survival Analysis
Quantifying and Rejecting Outliers: The Grubbs Test

