Related Experiment Video
Updated: Jan 21, 2026

Microarray-based Identification of Individual HERV Loci Expression: Application to Biomarker Discovery in Prostate Cancer
Published on: November 2, 2013
Identification of potential biomarkers on microarray data using distributed gene selection approach
Alok Kumar Shukla1, Diwakar Tripathi2
1Department of Information Technology, VNR Vignana Jyothi Institute of Engineering and Technology, Hyderabad, India.
This study introduces a novel two-stage feature selection (FS) method combining Spearman's Correlation and distributed filters. The approach enhances computational efficiency and classification accuracy for identifying biomarkers in gene expression data.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning in Genomics
Background:
- Gene expression datasets are crucial for cancer classification, but existing feature selection (FS) methods face limitations.
- Centralized data structures in traditional FS increase computational costs.
- Potential biomarkers may be overlooked by FS methods that do not consider feature interactions or distributed data.
Purpose of the Study:
- To develop a novel two-stage feature selection (FS) approach to address limitations of existing methods.
- To improve computational efficiency and classification accuracy for high-dimensional gene expression datasets.
- To identify highly discriminative genes for distinguishing sample classes.
Main Methods:
- Introduced a two-stage FS approach combining Spearman's Correlation (SC) and distributed filter methods.
- Employed vertical data distribution for distributed FS, followed by a merging procedure to update feature subsets.
- Quantified gene-gene and gene-class relationships to detect essential gene subsets.
Main Results:
- The proposed method demonstrated significantly improved computational time and classification accuracy compared to standard algorithms.
- Validated on six gene datasets using four classifiers: support vector machine, naïve Bayes, k-nearest neighbor, and decision tree.
- Outperformed traditional filter techniques including Relief-F, Information Gain, Minimum Redundancy Maximum Relevance, Joint Mutual Information, Chi-square, and t-test.
Conclusions:
- The novel two-stage FS approach effectively identifies discriminative genes from high-dimensional datasets.
- The method offers a more computationally efficient and accurate solution for biomarker discovery in gene expression analysis.
- Distributed FS combined with correlation analysis provides a robust framework for cancer classification and biomarker identification.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Data: Types and Distribution
Distributions in...
Selected Data About Geographic Locations
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...
Model Approaches for Pharmacokinetic Data: Physiological Models
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...

