Related Experiment Video
Updated: Jun 23, 2026

05:12
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
The properties of high-dimensional data spaces: implications for exploring gene and protein expression data
Robert Clarke1, Habtom W Ressom, Antai Wang
1Department of Oncology and Lombardi Comprehensive Cancer Center, Georgetown University School of Medicine, 3970 Reservoir Road NW, Washington, DC 20057, USA.
Nature Reviews. Cancer
|December 22, 2007
Summary
High-throughput genomic and proteomic technologies generate complex, high-dimensional data in cancer research. Understanding these data properties is crucial for accurate analysis and interpretation in translational science.
Area of Science:
- Genomics
- Proteomics
- Cancer Research
- Translational Science
Background:
- High-throughput genomic and proteomic technologies are integral to modern cancer research.
- These technologies generate vast amounts of high-dimensional data, where each sample has numerous measurements.
- Extracting meaningful biological and statistical information from this data is a significant challenge.
Purpose of the Study:
- To review the properties of high-dimensional data spaces encountered in genomic and proteomic studies.
- To discuss the challenges these properties pose for data analysis and interpretation.
- To provide insights relevant to translational science in cancer research.
Main Methods:
- Review of existing literature and methodologies in high-dimensional data analysis.
- Discussion of statistical and biological interpretation challenges.
- Focus on the context of cancer research and translational applications.
Main Results:
- High-dimensional data spaces possess unique properties that are often misunderstood or neglected.
- These properties can significantly impact the reliability of predictive models for diagnosis, prognosis, and therapy.
- Challenges include identifying key signaling networks and potential drug targets.
Conclusions:
- A thorough understanding of high-dimensional data properties is essential for effective analysis in cancer genomics and proteomics.
- Addressing these challenges is critical for advancing translational science and drug development.
- Improved data modeling and interpretation strategies are needed to fully leverage high-throughput technologies.
Related Concept Videos
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
Genome Size and the Evolution of New Genes
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.

