Related Experiment Video
Updated: Oct 1, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A new feature selection method based on feature distinguishing ability and network influence
Yanpeng Qi1, Benzhe Su1, Xiaohui Lin1
1School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, Liaoning, China.
Abstract:
The occurrence and development of diseases are related to the dysfunction of biomolecules (genes, metabolites, etc.) and the changes of molecule interactions. Identifying the key molecules related to the physiological and pathological changes of organisms from omics data is of great significance for disease diagnosis, early warning and drug-target prediction, etc. A novel feature selection algorithm based on the feature individual distinguishing ability and feature influence in the biological network (FS-DANI) is proposed for defining important biomolecules (features) to discriminate different disease conditions. The feature individual distinguishing ability is evaluated based on the overlapping area of the feature effective ranges in different classes. FS-DANI measures the feature network influence based on the module importance in the correlation network and the feature centrality in the modules. The feature comprehensive weight is obtained by combining the feature individual distinguishing ability and feature influence in the network. Then crucial feature subset is determined by the sequential forward search (SFS) on the feature list sorted according to the comprehensive weights of features. FS-DANI is compared with the six efficient feature selection methods on ten public omics datasets. The ablation experiment is also conducted. Experimental results show that FS-DANI is better than the compared algorithms in accuracy, sensitivity and specificity on the whole. On analyzing the gastric cancer miRNA expression data, FS-DANI identified two miRNAs (hsa-miR-18a* and hsa-miR-381), whose AUCs for distinguishing gastric cancer samples and normal samples are 0.959 and 0.879 in the discovery set and an independent validation set, respectively. Hence, evaluating biomolecules from the molecular level and network level is helpful for identifying the potential disease biomarkers of high performance.
Related Concept Videos
Outliers and Influential Points
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Frequency-dependent Selection
Types of Selection
Causes of Similarity-Dissimilarity Effect
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

