Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Sample Size Calculation01:19

Sample Size Calculation

6.7K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
6.7K
RNA-seq03:21

RNA-seq

12.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.0K
Regression Analysis01:11

Regression Analysis

8.4K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.4K
Regression Toward the Mean01:52

Regression Toward the Mean

7.0K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.0K
Microsoft Excel: Regression Analysis01:18

Microsoft Excel: Regression Analysis

1.6K
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
1.6K
The Binomial Theorem01:30

The Binomial Theorem

322
The Binomial Theorem is a foundational principle in algebra used to expand expressions raised to a power. It provides a structured approach for expanding binomials of the form (a+b)n, where a and b are variables or constants representing algebraic expressions, and n is a non-negative integer.The general form of the Binomial Theorem is:Each term in the expansion involves a binomial coefficient, which is calculated using factorials:The exponent of a in each term decreases from n to 0, while the...
322

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Standard knee radiographs enable deep learning inference of MRI-defined cartilage and meniscal damage in early knee osteoarthritis: a study using the osteoarthritis initiative database.

Frontiers in physiology·2026
Same author

A Phase 2 Feasibility Study Combining Pembrolizumab and Metformin in patients with Metastatic Head and Neck Cancer.

Clinical cancer research : an official journal of the American Association for Cancer Research·2026
Same author

Urolithin A activates aryl hydrocarbon receptor-NLRP6-mediated pathways in intestinal epithelial cells to modulate mucosal immunity and strengthen gut barrier integrity.

Nature communications·2026
Same author

Impact of high-fat Western diet on chronic lymphocytic leukemia disease progression and gut microbiome profile in Eμ-TCL1 mice.

bioRxiv : the preprint server for biology·2026
Same author

Amivantamab for Recurrent or Metastatic Adenoid Cystic Carcinoma: A Phase 2 Nonrandomized Clinical Trial.

JAMA otolaryngology-- head & neck surgery·2026
Same author

Characterization of the clonal hierarchy and immunophenotype of PTPN11 mutations in acute myeloid leukemia.

JCI insight·2026

Related Experiment Video

Updated: Jan 30, 2026

Characterization of In Vitro Differentiation of Human Primary Keratinocytes by RNA-Seq Analysis
07:29

Characterization of In Vitro Differentiation of Human Primary Keratinocytes by RNA-Seq Analysis

Published on: May 16, 2020

6.7K

Sample size calculations for the differential expression analysis of RNA-seq data using a negative binomial

Xiaohong Li1,2, Dongfeng Wu1, Nigel G F Cooper2

  • 1Department of Bioinformatics and Biostatistics, School of Public Health and Information Sciences, University of Louisville, Louisville, KY 40202, USA.

Statistical Applications in Genetics and Molecular Biology
|January 23, 2019
PubMed
Summary

We developed new sample size calculation methods for RNA sequencing (RNA-seq) studies using negative binomial regression. Our M1 method offers a computationally efficient approach for determining sample sizes in biomarker discovery experiments.

Keywords:
RNA-seqa Wald testdifferentially expressed genes (DEGs)power analysissample size

More Related Videos

Identification of Alternative Splicing and Polyadenylation in RNA-seq Data
08:35

Identification of Alternative Splicing and Polyadenylation in RNA-seq Data

Published on: June 24, 2021

6.4K
RNA-Seq Analysis of Differential Gene Expression in Electroporated Chick Embryonic Spinal Cord
11:13

RNA-Seq Analysis of Differential Gene Expression in Electroporated Chick Embryonic Spinal Cord

Published on: November 1, 2014

15.1K

Related Experiment Videos

Last Updated: Jan 30, 2026

Characterization of In Vitro Differentiation of Human Primary Keratinocytes by RNA-Seq Analysis
07:29

Characterization of In Vitro Differentiation of Human Primary Keratinocytes by RNA-Seq Analysis

Published on: May 16, 2020

6.7K
Identification of Alternative Splicing and Polyadenylation in RNA-seq Data
08:35

Identification of Alternative Splicing and Polyadenylation in RNA-seq Data

Published on: June 24, 2021

6.4K
RNA-Seq Analysis of Differential Gene Expression in Electroporated Chick Embryonic Spinal Cord
11:13

RNA-Seq Analysis of Differential Gene Expression in Electroporated Chick Embryonic Spinal Cord

Published on: November 1, 2014

15.1K

Area of Science:

  • Genomics and Bioinformatics
  • Statistical Genetics
  • Biomarker Discovery

Background:

  • High-throughput RNA sequencing (RNA-seq) is crucial for disease-related biomarker studies.
  • Negative binomial distribution is commonly used for RNA-seq read counts due to over-dispersion.
  • Accurate sample size estimation is vital for robust RNA-seq experimental design.

Purpose of the Study:

  • To propose two novel, explicit sample size calculation methods for RNA-seq data.
  • To utilize a negative binomial regression model incorporating dispersion parameters and size factors.
  • To provide a computationally efficient method for experimental design.

Main Methods:

  • Developed sample size formulas based on a negative binomial regression model.
  • Incorporated common dispersion parameter and size factor with a natural logarithm link function.
  • Employed a two-sided Wald test statistic for gene significance testing at FDR 0.05.

Main Results:

  • Proposed two new sample size calculation methods for RNA-seq studies.
  • Evaluated performance via simulation studies, comparing with existing methods.
  • Identified the M1 method as computationally efficient and suitable for quick estimation.

Conclusions:

  • The proposed methods provide explicit sample size calculations for RNA-seq experiments.
  • The M1 method is recommended for its computational efficiency in experimental design.
  • Demonstrated sample size estimation using real-world breast cancer RNA-seq data.