Related Experiment Video
Updated: Jun 24, 2026

05:12
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
High-dimensional data analysis: selection of variables, data compression and graphics--application to gene expression
Jürgen Läuter1, Friedemann Horn, Maciej Rosołowski
1Interdisciplinary Centre for Bioinformatics (IZBI), University of Leipzig, Härtelstr. 16-18, 04107 Leipzig, Germany. juergen.laeuter@med.ovgu.de
Biometrical Journal. Biometrische Zeitschrift
|April 10, 2009
Summary
This study introduces precise variable selection methods for high-dimensional data, like gene expression analysis. These techniques enhance statistical power and biological interpretation while preventing overfitting and controlling errors.
Area of Science:
- Bioinformatics
- Statistical Genetics
- Computational Biology
Background:
- High-dimensional data, such as gene expression data, presents significant challenges in statistical analysis.
- Effective variable selection is crucial for increasing statistical power and enabling biological interpretation in complex datasets.
- Existing methods may struggle with the scale and complexity of modern biological data, necessitating robust approaches.
Purpose of the Study:
- To present mathematically exact and effective procedures for variable selection in very high-dimensional settings.
- To enhance statistical power and facilitate biological interpretation through the construction of variable sets.
- To ensure statistical stability and avoid overfitting while maintaining control over familywise type I error.
Main Methods:
- Variables are selected by considering each single variable as the center of potential variable sets.
- Significance testing is performed using the Westfall-Young principle (resampling-based) or parametric spherical tests.
- A specific data compression technique is employed for generating graphical representations like heat maps and curves.
Main Results:
- The proposed procedures are effective for high-dimensional data, including gene expression analysis.
- High statistical power is achieved, and familywise type I error is controlled despite large data dimensions.
- Graphical representations using heat maps and curves are generated via a data compression technique.
Conclusions:
- The developed variable selection methods offer a robust solution for analyzing high-dimensional biological data.
- These procedures successfully balance statistical power, interpretability, and error control.
- The methodology is demonstrated effectively using gene expression data from B-cell lymphoma patients.
Related Concept Videos
Statgraphics
Statgraphics is a comprehensive statistical software suite designed for both basic and advanced data analysis. Originating in 1980 at Princeton University under Dr. Neil W. Polhemus, it was one of the pioneering tools for statistical computing on personal computers, with its public release in 1982 marking an early milestone in data science software. Over the years, it has evolved into a robust platform for data science, offering tools for regression analysis, ANOVA, multivariate statistics,...
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
Biostatistics: Overview
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...

