Related Experiment Videos
Implementing the Fisher's discriminant ratio in a k-means clustering algorithm for feature selection and data set
Thy-Hou Lin1, Huang-Te Li, Keng-Chang Tsai
1Institute of Molecular Medicine & Department of Life Science, National Tsing Hua University, Hsinchu, Taiwan 30013, ROC. thlin@life.nthu.edu.tw
Summary
This study introduces a novel method using Fisher's discriminant ratio and k-means clustering for feature selection and data trimming in HIV-1 protease inhibitors. The approach successfully identifies key topological descriptors, retaining 44% of inhibitors with enhanced class sensitivity.
Area of Science:
- Computational Chemistry
- Cheminformatics
- Bioinformatics
Background:
- Accurate feature selection and data trimming are crucial for understanding structure-activity relationships in drug discovery.
- HIV-1 protease inhibitors represent a significant class of antiviral drugs, necessitating efficient analysis methods.
- Existing methods may not optimally balance feature selection with outlier identification for complex datasets.
Purpose of the Study:
- To develop and validate a simultaneous feature selection and data trimming method for HIV-1 protease inhibitors.
- To identify class-sensitive molecular descriptors that correlate with inhibitory activity (pKi).
- To improve the quality of datasets for further quantitative structure-activity relationship (QSAR) analyses.
Main Methods:
- Implementation of Fisher's discriminant ratio as a class separability criterion within a k-means clustering algorithm.
- Application of feature evaluation indices including Shannon entropy, linear regression, and stepwise variable selection for descriptor filtering.
- Division of 221 HIV-1 protease inhibitors into five classes to optimize feature selection.
Main Results:
- The method successfully identified topological descriptors that are well-correlated with pKi values.
- K-means clustering effectively identified and removed outliers (trimmed inhibitors).
- Retaining 44% (98 inhibitors) with three selected descriptors resulted in significantly more class-sensitive data, validated by PLS regression.
Conclusions:
- The combined Fisher's discriminant ratio and k-means clustering approach offers an effective strategy for simultaneous feature selection and data trimming.
- Selected topological descriptors provide valuable insights into the activity of HIV-1 protease inhibitors.
- The refined dataset enhances statistical significance for subsequent QSAR modeling and drug design.