A systematic comparative evaluation of biclustering techniques
Victor A Padilha1, Ricardo J G B Campello2,3
1Institute of Mathematics and Computer Sciences, University of São Paulo, São Carlos, SP, Brazil. victorpadilha@usp.br.
BMC Bioinformatics
|January 25, 2017
Summary
This study compares seventeen biclustering algorithms on diverse datasets. No single algorithm excels at all tasks; the best choice depends on specific data patterns and research goals.
Area of Science:
- Bioinformatics
- Computational Biology
- Data Mining
Background:
- Biclustering simultaneously analyzes rows and columns of data matrices.
- It is widely used for gene expression data analysis, identifying genes in multiple pathways active under specific conditions.
- Previous comparative studies used limited algorithms and datasets with suboptimal evaluation metrics.
Purpose of the Study:
- To conduct a comprehensive comparison of biclustering algorithms.
- To evaluate algorithm performance using more appropriate metrics.
- To provide guidance for selecting biclustering methods based on specific application needs.
Main Methods:
- Seventeen biclustering algorithms were evaluated.
- Experiments were performed on three synthetic and two real data collections, using a larger number of datasets.
- Synthetic data experiments covered noise levels, bicluster numbers, overlap (symmetric and asymmetric), and sizes.
- Real data experiments utilized gene set enrichment and clustering accuracy for assessment.
Main Results:
- Each evaluated biclustering algorithm demonstrated strengths in specific tasks.
- Algorithm performance varied significantly across different experimental scenarios and data types.
- No single algorithm universally outperformed others across all tested conditions.
Conclusions:
- The optimal biclustering algorithm is task-dependent.
- Selection should consider the specific patterns to be detected and the nature of the data.
- This study offers a foundation for informed algorithm selection in biclustering applications.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
7.1K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.1K
Multiple Comparison Tests
4.5K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.5K
Cluster Sampling Method
15.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.3K


