Related Experiment Video
Updated: Nov 4, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
The Unsupervised Feature Selection Algorithms Based on Standard Deviation and Cosine Similarity for Genomic Data
Juanying Xie1, Mingzhao Wang1,2, Shengquan Xu2
1School of Computer Science, Shaanxi Normal University, Xi'an, China.
This study introduces unsupervised feature selection methods (SCFS, SCEFS, SCRFS, SCAFS) to analyze high-dimensional genomic data. These methods effectively identify cancer biomarkers for improved diagnostics and pathology research.
Area of Science:
- Genomics
- Bioinformatics
- Machine Learning
Background:
- Genomic data analysis faces challenges due to high dimensionality, limited samples, and class imbalance.
- Unsupervised feature selection is crucial for identifying relevant biological markers without prior labels.
Purpose of the Study:
- To propose novel unsupervised feature selection techniques for genomic data.
- To address the challenges of high dimensionality and class imbalance in cancer genomics.
- To identify robust biomarkers for cancer classification and research.
Main Methods:
- Developed Standard deviation and Cosine similarity based Feature Selection (SCFS) using discernibility and independence.
- Derived three algorithms: SCEFS, SCRFS, and SCAFS, based on different cosine similarity measures.
- Utilized KNN and SVM classifiers with selected features on 18 cancer genomic datasets.
Main Results:
- The proposed algorithms (SCEFS, SCRFS, SCAFS) successfully identified stable biomarkers with strong classification capabilities across 18 cancer datasets.
- Feature importance was determined by a 2D space plotting discernibility against independence.
- Functional analysis linked identified biomarkers to gene regulation levels relevant to cancer occurrence.
Conclusions:
- The SCFS-derived algorithms offer a powerful approach for unsupervised feature selection in high-dimensional genomic data.
- Identified biomarkers provide insights into cancer pathology, aiding drug development and early diagnosis.
- This method has significant implications for advancing cancer research and clinical applications.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Standard Deviation of Calculated Results
A broad Gaussian distribution curve has a wider standard deviation, representing a data set with...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Standard Deviation

