Related Experiment Video
Updated: May 30, 2025

08:49
Improving Small RNA-seq: Less Bias and Better Detection of 2'-O-Methyl RNAs
Published on: September 16, 2019
7.6K
Correcting scale distortion in RNA sequencing data
Christopher Thron1, Farhad Jafari2
1Department of Science and Mathematics, Texas A &M University-Central Texas, Killeen, TX, 76549, USA. thron@tamuct.edu.
BMC Bioinformatics
|January 28, 2025
Summary
Researchers developed new methods to correct expression-level biases in RNA sequencing (RNA-seq) data. These techniques improve the accuracy of gene expression analysis for disease research.
Area of Science:
- Genomics and Bioinformatics
- Molecular Biology
- Computational Biology
Background:
- RNA sequencing (RNA-seq) is a standard method for measuring gene expression across entire genomes.
- Accurate gene expression data is crucial for identifying genetic factors in diseases through population studies.
- Existing normalization methods may not fully correct for all sources of error in RNA-seq data.
Purpose of the Study:
- To identify and correct expression level-dependent biases in RNA sequencing data that persist after standard normalization.
- To improve the accuracy of gene-gene correlation estimations and statistical tests in population studies.
- To enhance the reliability of RNA-seq data for clinical and research applications.
Main Methods:
- Analysis of multiple RNA-seq datasets from TCGA, SU2C, and GTEx databases.
- Application of local averaging to detect sample-specific, expression-dependent biases.
- Development and application of two novel nonlinear transforms to correct identified biases.
- Utilized a new simulation methodology to assess the impact of corrections on statistical tests.
Main Results:
- Expression level-dependent biases were detected in all analyzed RNA-seq datasets, varying between samples.
- These biases were shown to negatively impact gene-gene correlation and differential expression analyses.
- The proposed nonlinear transforms effectively removed per-sample biases, reduced variance, and improved correlation distributions.
- Data correction led to a 3-5% improvement in sensitivity and specificity for two-population tests.
Conclusions:
- Novel nonlinear transforms can accurately correct for expression level-dependent biases in RNA-seq data.
- Bias correction enhances the reliability of gene-gene relationship analysis and statistical power in population studies.
- These findings offer improved methods for utilizing clinical RNA-seq data, potentially leading to new discoveries.

