Related Experiment Video
Updated: May 4, 2026

13:19
Microarray-based Identification of Individual HERV Loci Expression: Application to Biomarker Discovery in Prostate Cancer
Published on: November 2, 2013
19.4K
Proximity measures for clustering gene expression microarray data: a validation methodology and a comparative
Pablo A Jaskowiak1, Ricardo J G B Campello1, Ivan G Costa2
1University of São Paulo, São Carlos.
IEEE/ACM Transactions on Computational Biology and Bioinformatics
|December 17, 2013
Summary
Choosing the right proximity measure is crucial for gene expression microarray data clustering. This study reveals that less common measures often outperform standard ones like Pearson, highlighting the need for scenario-specific selection in gene expression analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Cluster analysis is a primary step for extracting insights from gene expression microarray data.
- Selecting an appropriate proximity measure is critical for effective clustering, yet guidelines are lacking.
- Pearson is widely used, but the performance of other measures remains largely unexamined.
Purpose of the Study:
- To investigate the impact of different proximity measures on microarray data clustering.
- To evaluate the performance of 16 proximity measures across diverse gene expression datasets.
- To provide guidance on selecting proximity measures for specific experimental contexts.
Main Methods:
- Evaluated 16 proximity measures on 52 gene expression microarray datasets from cancer and time-course experiments.
- Developed a benchmark and a novel methodology, Intrinsic Biological Separation Ability (IBSA), for time-course data evaluation.
- Compared the performance of commonly used measures (e.g., Pearson, Spearman, Euclidean distance) against less explored alternatives.
Main Results:
- Measures not frequently used in gene expression literature demonstrated superior performance compared to standard measures.
- The optimal proximity measure varied significantly between time-course and cancer gene expression data.
- The IBSA methodology and benchmark offer a standardized approach for future research on time-course data.
Conclusions:
- The choice of proximity measure for clustering gene expression data should be tailored to the specific experimental scenario (e.g., time-course vs. cancer).
- Rarely used proximity measures can yield better clustering results than commonly employed ones.
- The developed IBSA methodology and benchmark are valuable resources for advancing research in gene time-course data analysis.

