Related Experiment Video
Updated: May 31, 2026

08:35
Identification of Alternative Splicing and Polyadenylation in RNA-seq Data
Published on: June 24, 2021
Identifying Relevant Covariates in RNA-seq Analysis by Pseudo-Variable Augmentation
1Department of Mathematics and Statistics, Old Dominion University, Norfolk, VA 23529 USA.
Summary
This study introduces a novel variable selection method for RNA-sequencing (RNA-seq) data. The method accurately identifies relevant covariates, improving the detection of differentially expressed genes while controlling false selections.
Area of Science:
- Genomics
- Bioinformatics
- Statistical Genetics
Background:
- RNA-sequencing (RNA-seq) is crucial for identifying differentially expressed genes.
- RNA-seq datasets often contain relevant and irrelevant covariates that can complicate gene expression analysis.
- Ignoring or incorrectly adjusting for covariates can compromise the accurate identification of differentially expressed genes.
Purpose of the Study:
- To develop a robust variable selection method for RNA-sequencing data.
- To accurately identify relevant covariates while controlling the false selection rate.
- To improve the detection of differentially expressed genes in the presence of complex covariate structures.
Main Methods:
- Proposed a novel variable selection method utilizing pseudo-variables.
- Implemented a strategy to control the expected proportion of irrelevant selected covariates.
- Developed the FSRAnalysisBS function within the R package csrnaseq.
Main Results:
- The proposed method accurately selects relevant covariates.
- The method effectively controls the false selection rate below a specified threshold.
- Demonstrated superior performance compared to existing methods for differential gene expression analysis with covariates.
Conclusions:
- The new variable selection approach enhances the reliability of differential gene expression analysis.
- This method offers a significant improvement for RNA-seq studies with complex covariate data.
- The csrnaseq R package provides a practical implementation for researchers.
Related Concept Videos
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Ribosome Profiling
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
RACE - Rapid Amplification of cDNA Ends
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific primer.
Since the...
Since the...

