Related Experiment Video
Updated: Jul 18, 2025

Analyzing Multifactorial RNA-Seq Experiments with DiCoExpress
Published on: July 29, 2022
Categorical Data Analysis for High-Dimensional Sparse Gene Expression Data
Niloufar Dousti Mousavi1, Hani Aldirawi2, Jie Yang1
1Department of Mathematics, Statistics, and Computer Science, University of Illinois at Chicago, Chicago, IL 60607, USA.
We developed a statistical method for analyzing high-dimensional omics data. This approach effectively identifies key genes for cancer tumor classification and prognostic signatures.
Area of Science:
- Biostatistics
- Bioinformatics
- Genomics
Background:
- Analyzing categorical data with high-dimensional sparse covariates, common in omics data, presents significant challenges.
- Existing statistical methods may not adequately address the complexities of variable and model selection in such datasets.
Purpose of the Study:
- To introduce a comprehensive statistical procedure for categorical data analysis in the context of high-dimensional omics data.
- To enable variable screening, model selection, response category ordering, and variable selection for complex biological datasets.
Main Methods:
- A statistical procedure based on multinomial logistic regression analysis was developed.
- The procedure incorporates variable screening, model selection, order selection for response categories, and variable selection.
- The method was applied to high-dimensional gene expression data from 801 patients across five cancerous tumor types.
Main Results:
- A finalized model with 74 genes demonstrated extremely low cross-entropy loss and zero predictive error rate via five-fold cross-validation.
- Two additional models, comprising 31 and 4 genes, were identified as potential prognostic multi-gene signatures.
Conclusions:
- The proposed statistical procedure is effective for categorical data analysis with high-dimensional sparse covariates, particularly in omics research.
- The identified gene signatures offer potential for improved cancer tumor classification and prognosis.
More Related Videos
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Friedman Two-way Analysis of Variance by Ranks
Comparing the Survival Analysis of Two or More Groups
Quantifying and Rejecting Outliers: The Grubbs Test
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...

