Related Experiment Video
Updated: Nov 23, 2025

06:17
Analysis of Multidimensional Microscopy Data Using Cell-ACDC
Published on: November 7, 2025
169
On linear dimension reduction based on diagonalization of scatter matrices for bioinformatics downstream analyses.
Daniel Fischer1, Klaus Nordhausen2, Hannu Oja3
1Natural Resources Institute Finland (Luke), Applied Statistical Methods, Myllytie 1, 31600 Jokionen, Finland.
Heliyon
|January 1, 2021
Summary
This study explores dimension reduction techniques for analyzing small sample size, large variable datasets common in bioinformatics. It demonstrates their utility in identifying significant biological patterns, using a prostate cancer microarray dataset as an example.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Learning
Background:
- Dimension reduction is crucial for high-dimensional data analysis.
- Classical methods like PCA, ICA, and SIR are often formulated using scatter matrices.
- Existing theoretical guarantees typically require sample size to exceed the number of variables, limiting applicability in bioinformatics.
Purpose of the Study:
- To investigate the applicability of dimension reduction tools for small-n-large-p datasets.
- To explore the use of scatter matrix functionals for identifying relevant subspaces.
- To assess the effectiveness of these methods in bioinformatics contexts.
Main Methods:
- Review of classical supervised and unsupervised dimension reduction techniques.
- Formulation of methods using scatter matrix functionals as measures of multivariate dispersion.
- Application and illustration using a prostate cancer microarray dataset.
Main Results:
- Dimension reduction tools can be adapted for small-n-large-p data.
- Scatter matrices highlight different data features and reveal structures when compared.
- The study successfully illustrates subspace identification on a real-world bioinformatics dataset.
Conclusions:
- Dimension reduction methods, particularly those based on scatter matrices, are valuable for analyzing high-dimensional biological data.
- These techniques can effectively identify relevant subspaces even when the sample size is small relative to the number of variables.
- The findings support the use of these methods in bioinformatics for uncovering biological insights.

