Related Experiment Videos
Biclustering of gene expression data by Non-smooth Non-negative Matrix Factorization
Pedro Carmona-Saez1, Roberto D Pascual-Marqui, F Tirado
1BioComputing Unit, National Center of Biotechnology, Campus Universidad Autónoma de Madrid, 28049. Spain. pcarmona@cnb.uam.es
BMC Bioinformatics
|March 1, 2006
Summary
This study introduces Non-smooth Non-Negative Matrix Factorization (nsNMF) to find localized gene expression patterns in large datasets. The method effectively clusters genes and conditions, revealing biological insights into physiological states.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Microarray technologies generate large gene expression datasets across numerous experimental conditions.
- Analyzing these datasets to find local structures of co-expressed genes is a significant challenge.
- Such structures can offer insights into biological processes linked to different physiological states.
Purpose of the Study:
- To develop a methodology for clustering genes and conditions within specific data subsets.
- To identify localized patterns in large-scale gene expression data.
Main Methods:
- Utilized a novel data mining technique: Non-smooth Non-Negative Matrix Factorization (nsNMF).
- Applied nsNMF to cluster genes and conditions exhibiting related expression patterns in data sub-portions.
- Validated the methodology on synthetic and large, heterogeneous gene expression datasets.
Main Results:
- The nsNMF approach successfully identified localized features in gene expression data.
- Consistent expression patterns were found across subsets of experimental conditions for specific gene sets.
- Uncovered structures demonstrated biological relevance, linking gene functions to experimental conditions' phenotypes.
Conclusions:
- The proposed nsNMF methodology is a valuable tool for analyzing large, heterogeneous gene expression datasets.
- It effectively identifies complex gene-condition relationships often missed by standard clustering algorithms.