Related Experiment Video
Updated: May 13, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Double weighted k nearest neighbours for binary classification of high dimensional genomic data
Amjad Ali1, Zardad Khan2, Hailiang Du3,4
1Department of Statistics and Bussines Analytics, United Arab Emirates University, Al Ain, United Arab Emirates.
Abstract:
High dimensional gene expression datasets consist of a large number of genes, many of which do not play a significant role in classifying tissue samples. The high dimensional nature of this type of data, characterized by a large number of gene features substantially exceeding its sample size, makes it challenging for existing methods to work efficiently in terms of prediction accuracy and execution time. To address this issue, a new classification procedure called double weighted k nearest neighbours ([Formula: see text]) is proposed. [Formula: see text] is specifically designed for gene expression data and incorporates feature weights derived from genes' ability to express deferentially between classes. Features weights are derived in a manner that automatically increase the impact of informative features while decreasing it for features that are less/non informative. To achieve this goal, the estimated weighted distances from the observations in the k nearest neighbourhood to the test point are used in an exponential function. The outputs of the function are summed for both the classes separately and the test point is assigned the class label with the largest sum. By utilizing the proposed weighting method based on the differential capability of genes, the [Formula: see text] method aims to achieve robust and efficient classification by allowing only the most informative features/genes to contribute to the classification task. Experimental evaluations, in comparison with several methods, i.e., standard [Formula: see text], weighted k nearest neighbours classifier ([Formula: see text]), random k nearest neighbour ([Formula: see text]), extended neighbourhood rule ensemble (ExNRule), k conditional nearest neighbour ([Formula: see text]), [Formula: see text] ensemble and support vector machines (SVM), demonstrate the effectiveness of [Formula: see text] in accurately classifying gene expression datasets. Overall, [Formula: see text] presents a promising approach for gene expression data analysis through the two fold weighted distance calculation strategy using classification accuracy, Cohen's kappa, sensitivity and [Formula: see text]score as performance metrics.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Karyotyping

