Investigating the minimum required number of genes for the classification of neuromuscular disease microarray data
Argiris Sakellariou1, Despina Sanoudou, George Spyrou
1Biomedical Research Foundation, Academy of Athens, and the Department of Informatics and Telecommunications, National and Kapodistrian University of Athens, Athens 115 27, Greece. asake@bioacademy.gr
Abstract:
The discovery of potential microarray markers, which will expedite molecular diagnosis/prognosis and provide reliable results to clinical decision-making and treatment selection for patients, is of paramount importance. Feature selection techniques, which aim at minimizing the dimensionality of the microarray data by keeping the most statistically significant genes, are a powerful approach toward this goal. In this paper, we investigate the minimum required subsets of genes, which best classify neuromuscular disease data. For this purpose, we implemented a methodology pipeline that facilitated the use of multiple feature selection methods and subsequent performance of data classification. Five feature selection methods on datasets from ten different neuromuscular diseases were utilized. Our findings reveal subsets of very small number of genes, which can successfully classify normal/disease samples. Interestingly, we observe that similar classification results may be obtained from different subsets of genes. The proposed methodology can expedite the identification of small gene subsets with high-classification accuracy that could ultimately be used in the genetics clinics for diagnostic, prognostic, and pharmacogenomic purposes.
More Related Videos
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
08:22A Novel Strategy Combining Array-CGH, Whole-exome Sequencing and In Utero Electroporation in Rodents to Identify Causative Genes for Brain Malformations
Published on: December 1, 2017
