Related Experiment Video
Updated: Aug 9, 2026

13:33
Infinium Assay for Large-scale SNP Genotyping Applications
Published on: November 19, 2013
A generalized clustering problem, with application to DNA microarrays
1New York University School of Medicine, Division of Biostatistics, USA. ilana.belitskaya@med.nyu.edu
Summary
This study introduces a novel algorithm for discovering multiple clustering structures within data, moving beyond traditional single-structure assumptions. The method clusters both variables and observations, with applications in gene expression analysis.
Area of Science:
- Data Science
- Bioinformatics
- Statistics
Background:
- Traditional cluster analysis assumes a single, unique clustering structure for observations.
- This assumption is limiting as data can exhibit multiple valid clusterings based on different variable subsets.
Purpose of the Study:
- To develop and present an algorithm for identifying multiple, distinct clustering structures within a dataset.
- To generalize cluster analysis beyond the assumption of a unique underlying structure.
Main Methods:
- Proposes a novel algorithm that simultaneously clusters variables and observations.
- Variable clustering uses a dissimilarity measure based on nearest-neighbor graphs.
- Observation clustering employs weighted distances, with weights derived from variable clusters.
Main Results:
- The algorithm can identify multiple clustering structures.
- The number of identified clustering structures is determined by the number of variable clusters.
- Demonstrates applicability to gene expression data analysis.
Conclusions:
- The proposed method effectively estimates multiple clustering structures, offering a more comprehensive data analysis approach.
- This generalized clustering framework is particularly relevant for complex datasets like gene expression data.
Related Concept Videos
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Applications of Molecular Taxonomy
Molecular taxonomy has revolutionized the understanding and classification of bacteria, providing precise insights into their diversity, evolutionary relationships, and ecological roles. By utilizing molecular techniques such as DNA sequencing and fingerprinting, researchers have made significant strides in various fields related to bacterial studies.Resolving Taxonomic AmbiguitiesMolecular taxonomy has been instrumental in distinguishing closely related bacterial species initially thought to...
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
