Related Experiment Video
Updated: Jul 12, 2026

07:28
JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
An improved algorithm for clustering gene expression data
Sanghamitra Bandyopadhyay1, Anirban Mukhopadhyay, Ujjwal Maulik
1Machine Intelligence Unit, Indian Statistical Institute, Kolkata-700108, India.
Bioinformatics (Oxford, England)
|August 28, 2007
Summary
A novel two-stage clustering algorithm enhances gene expression analysis by handling data uncertainty. This method, utilizing genetic algorithms and fuzzy C-means, outperforms existing techniques for biological data interpretation.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Microarray technology enables simultaneous monitoring of numerous gene expression levels over time.
- Analyzing this high-dimensional data requires robust clustering methods to address inherent uncertainty, noise, and imprecision.
Purpose of the Study:
- To propose a novel two-stage clustering algorithm for analyzing gene expression data.
- To address the challenge of data uncertainty and imprecision in microarray datasets.
Main Methods:
- A two-stage clustering approach combining a variable string length genetic scheme and a multiobjective genetic clustering algorithm.
- Incorporation of the concept of significant membership of points to multiple classes.
- Utilizing an iterated version of Fuzzy C-Means for clustering.
Main Results:
- The proposed two-stage algorithm demonstrated significant superiority over established methods like average linkage, Self-Organizing Map (SOM), and weighted Chinese Restaurant-based clustering (CRC).
- Validation was performed on both artificial and real-life gene expression datasets.
- The biological relevance of the obtained clustering solutions was assessed.
Conclusions:
- The developed two-stage clustering algorithm offers a more effective approach for gene expression data analysis.
- The method's ability to handle data uncertainty and its superior performance highlight its potential for biological discovery.
Related Concept Videos
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Cluster Sampling Method
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Extraction: Advanced Methods
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is formed in...
