Related Experiment Video
Updated: Mar 5, 2026

09:45
Measuring mRNA Levels Over Time During the Yeast S. cerevisiae Hypoxic Response
Published on: August 10, 2017
8.7K
Variance component score test for time-course gene set analysis of longitudinal RNA-seq data
Denis Agniel1, Boris P Hejblum2
1Department of Biomedical Informatics, Harvard Medical School, 10 Shattuck St, Boston, MA 02115, USA.
Biostatistics (Oxford, England)
|March 24, 2017
Summary
We developed tcgsaseq, a new statistical method for analyzing RNA-seq gene expression changes over time. This model-free approach offers improved stability and power for detecting longitudinal gene set variations compared to existing methods.
Area of Science:
- Genomics
- Bioinformatics
- Statistical Genetics
Background:
- RNA-sequencing (RNA-seq) is replacing microarrays for gene expression measurement.
- RNA-seq data, measured as counts, require adapted statistical tools due to inherent heteroscedasticity.
- Nonparametric regression has been proposed to model RNA-seq counts as continuous variables.
Purpose of the Study:
- To propose tcgsaseq, a principled, model-free, and efficient method for detecting longitudinal changes in predefined RNA-seq gene sets.
- To identify gene sets with expression that varies over time using a novel variance component score test.
Main Methods:
- tcgsaseq employs a variance component score test that accounts for covariates and heteroscedasticity without assuming parametric distributions for transformed counts.
- The method is computationally efficient, featuring a simple test statistic form and limiting distribution.
- A permutation-based version is available for small sample sizes.
Main Results:
- tcgsaseq demonstrates excellent statistical properties, including enhanced stability and power compared to ROAST, edgeR, and DESeq2.
- Existing methods like ROAST, edgeR, and DESeq2 may fail to control type I error rates in certain realistic scenarios.
- The method was validated using simulated data and two real biological datasets.
Conclusions:
- tcgsaseq provides a robust and powerful tool for analyzing longitudinal RNA-seq data.
- The proposed method addresses limitations of current state-of-the-art techniques in gene set analysis.
- The tcgsaseq method is publicly available as an R package for the research community.
Related Concept Videos
Variability: Analysis
581
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
581
Variance
13.0K
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.
The standard deviation measures the spread in the same units as the data....
The standard deviation measures the spread in the same units as the data....
13.0K
Coefficient of Variation
8.9K
The coefficient of variation measures the dispersion of the data points or distribution around the mean. Using the coefficient of variation, we can compare two data series with drastically different means or different units of measurement. The coefficient of variation for a sample and a population is expressed as a percentage of the ratio of standard deviation to the mean.
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...
8.9K
Comparing Copy Number Variations and SNPs
18.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.9K
Variation
8.2K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
8.2K
Significance Testing: Overview
12.9K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
12.9K

