Related Experiment Video
Updated: Mar 1, 2026

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
1.4K
Subspace Weighting Co-Clustering of Gene Expression Data
Summary
This study introduces Subspace Weighting Co-Clustering (SWCC), a novel algorithm for analyzing high-dimensional gene expression data. SWCC effectively identifies gene contributions to sample clusters, improving gene clustering and selection for enhanced classification performance.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Microarray technology generates extensive gene expression data.
- Clustering algorithms are vital for exploring this data.
- Co-clustering is beneficial as genes may correlate with sample subsets.
Purpose of the Study:
- To propose a novel co-clustering algorithm, Subspace Weighting Co-Clustering (SWCC), for high-dimensional gene expression data.
- To introduce a gene subspace weight matrix to identify gene contributions to sample clusters.
- To develop an iterative algorithm for solving the co-clustering objective function.
Main Methods:
- Developed the Subspace Weighting Co-Clustering (SWCC) algorithm.
- Introduced a gene subspace weight matrix into the co-clustering objective function.
- Designed an iterative algorithm to compute the subspace weight matrix during co-clustering.
Main Results:
- SWCC demonstrated encouraging performance compared to six state-of-the-art clustering algorithms on ten gene expression datasets.
- The algorithm was successfully applied to gene clustering and selection.
- Selected genes using SWCC improved the classification performance of Random Forests.
Conclusions:
- SWCC is an effective novel algorithm for co-clustering high-dimensional gene expression data.
- The gene subspace weight matrix is crucial for identifying gene contributions.
- SWCC facilitates improved gene selection for enhanced biological data analysis and machine learning applications.
Related Concept Videos
Cluster Sampling Method
15.2K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.2K
Weighted Mean
7.1K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
7.1K
DNA Microarrays
21.5K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
21.5K
Extraction: Partition and Distribution Coefficients
5.2K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
5.2K
Cell Specific Gene Expression
5.7K
5.7K
Cell Specific Gene Expression
16.7K
Multicellular organisms contain a variety of structurally and functionally distinct cell types, but the DNA in all the cells originated from the same parent cells. The differences in the cells can be attributed to the differential gene expression. Liver cells, whose functions include detoxification of blood, production of bile to metabolize fats, and synthesis of proteins essential for metabolism, must express a specific set of genes to perform their functions. Gene expression also varies with...
16.7K

