Related Experiment Video
Updated: Feb 13, 2026

Identification of Alternative Splicing and Polyadenylation in RNA-seq Data
Published on: June 24, 2021
Assessment of data transformations for model-based clustering of RNA-Seq data
Janelle R Noel-MacDonnell1,2, Joseph Usset1, Ellen L Goode3
1Department of Biostatistics, University of Kansas Medical Center, Kansas City, KS, United States of America.
Data transformations can improve clustering accuracy for RNA-Seq (RNA sequencing) data. Applying transformations to make RNA-Seq data more Gaussian-like enhanced clustering performance and accurate identification of gene clusters.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- RNA-Seq data analysis differs significantly from microarray analysis.
- RNA-Seq data often exhibit non-normal distributions (Poisson or negative binomial), unlike microarray data which are typically normal.
- Limited research exists on the impact of data transformations on Gaussian model-based clustering for RNA-Seq data.
Purpose of the Study:
- To investigate the impact of different data transformations on Gaussian model-based clustering performance for RNA-Seq data.
- To assess how these transformations affect the accuracy of estimating the correct number of clusters.
- To evaluate clustering performance using metrics like adjusted rand index, clustering error rate, and concordance index.
Main Methods:
- Simulated RNA-Seq data were generated with varying cluster sizes and numbers.
- Four data transformations were applied: naïve, logarithmic, Blom, and variance stabilizing transformation.
- Gaussian model-based clustering was performed on the transformed data.
Main Results:
- Data transformations generally improved clustering performance compared to no transformation.
- Transformations that made the data distribution more Gaussian-like led to better clustering outcomes.
- The choice of transformation impacted the accuracy in estimating the number of clusters.
Conclusions:
- Data transformations are beneficial for applying Gaussian model-based clustering to RNA-Seq data.
- Selecting appropriate transformations can enhance clustering accuracy and the reliability of cluster number estimation.
- Further research is warranted to optimize transformation strategies for diverse RNA-Seq datasets.
Related Concept Videos
Assessment of the Gastrointestinal System I: Subjective Data
Health History
The initial step in assessing the GI system is obtaining a comprehensive health history. This includes inquiring about the patient's history or presence of problems...
Assessment of the Cardiovascular System I: Subjective Data
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Data Reporting and Recording
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...

