Related Experiment Videos
An information theoretic approach for analyzing temporal patterns of gene expression.
Jyotsna Kasturi1, Raj Acharya, Murali Ramanathan
1Department of Computer Science and Engineering, Pennsylvania State University, University Park, PA 16802, USA. jkasturi@cse.psu.edu
Bioinformatics (Oxford, England)
|March 4, 2003
Summary
Information theoretic methods, like Kullback-Leibler divergence, effectively identify temporal gene expression patterns from microarray data. This approach outperforms traditional clustering methods for analyzing complex biological datasets.
Area of Science:
- Bioinformatics
- Computational Biology
- Systems Biology
Background:
- Microarray technology enables simultaneous measurement of thousands of mRNA expression levels.
- Gene expression data is rich but requires advanced mining for biological insights.
- Information theoretic methods can quantify similarities and dissimilarities in data distributions.
Purpose of the Study:
- To investigate information theoretic data mining approaches for discovering temporal gene expression patterns.
- To apply these methods to array-derived gene expression data.
Main Methods:
- Utilized Kullback-Leibler (KL) divergence, an information-theoretic measure of distribution dissimilarity.
- Employed an unsupervised self-organizing map (SOM) algorithm in conjunction with KL divergence.
- Analyzed two published, array-derived gene expression datasets.
Main Results:
- KL divergence-based clustering identified superior temporal patterns compared to hierarchical clustering with Pearson correlation.
- The discovered patterns demonstrated biological significance upon examination.
- The KL clustering method proved more effective in uncovering meaningful temporal trends in gene expression.
Conclusions:
- Information theoretic methods, specifically KL divergence with SOM, are powerful tools for analyzing temporal gene expression patterns.
- This approach offers advantages over traditional correlation-based clustering methods for microarray data.
- The findings highlight the utility of advanced computational techniques in extracting biological knowledge from high-throughput experiments.