Related Experiment Video
Updated: Jul 18, 2026

13:14
Global Gene Expression Analysis Using a Zebrafish Oligonucleotide Microarray Platform
Published on: August 10, 2009
Metric for measuring the effectiveness of clustering of DNA microarray expression
Raja Loganantharaj1, Satish Cheepala, John Clifford
1Bioinformatics Research Lab, University of Louisiana at Lafayette, PO Box 44330, Lafayette, LA 70504, USA. logan@cacs.louisiana.edu
BMC Bioinformatics
|November 23, 2006
Summary
A new metric evaluates gene clustering effectiveness using biological features, aiding algorithm selection. This method measures both within-cluster similarity and between-cluster distinctiveness for improved gene expression analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Microarray technology enables simultaneous expression analysis of thousands of genes.
- Gene expression data is often clustered using algorithms like hierarchical and k-means clustering with metrics such as Euclidean distance or Pearson correlation.
- Current validation indices for gene clustering lack direct correlation with biological features, confusing practitioners.
Purpose of the Study:
- To propose a novel metric for evaluating gene clustering algorithm effectiveness.
- To assess clustering based on biological features like biological processes and molecular functions.
- To provide a quantitative measure for comparing different clustering algorithms.
Main Methods:
- Developed a metric quantifying inter-cluster cohesiveness and intra-cluster separation.
- Applied the metric to gene expression data from a study on retinoids and cancer suppression.
- Evaluated hierarchical and k-means clustering with Euclidean and Pearson correlation distances.
Main Results:
- Genes with similar expression profiles are more closely related to biological processes than molecular functions.
- The proposed metric effectively measures both inter- and intra-cluster cohesiveness.
- The metric provides a single quantitative value (0 indicates maximum cohesiveness and separation) for easier algorithm comparison.
Conclusions:
- The best gene clustering algorithms should exhibit biological feature-based cohesiveness within clusters and maximal separation between clusters.
- The proposed metric is novel, easy to compute, and offers a unified measure of clustering effectiveness.
- The metric can be extended to other gene features (e.g., DNA binding sites, protein interactions) and applied to other domains.

