Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Systematic Error: Methodological and Sampling Errors01:15

Systematic Error: Methodological and Sampling Errors

In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Identifying Statistically Significant Differences: The F-Test01:14

Identifying Statistically Significant Differences: The F-Test

The F-test is used to compare two sample variances to each other or compare the sample variance to the population variance. It is used to decide whether an indeterminate error can explain the difference in their values. The underlying assumptions that allow the use of the F-test include the data set or sets are normally distributed, and the data sets are independent of each other. The test statistic F is calculated by dividing one variance by another. In other words, the square of one standard...
One-Way ANOVA: Equal Sample Sizes01:15

One-Way ANOVA: Equal Sample Sizes

One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Random Error01:04

Random Error

Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
Types of Errors: Detection and Minimization01:12

Types of Errors: Detection and Minimization

Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Low-intensity stimulation drives macrophage efferocytosis via ACSL4 lipid remodeling and CCL9-CCR1 signaling for tendon-bone healing.

Science advances·2026
Same author

Tissue-aware elastic net decomposition reveals shared and lineage-specific drug response biomarkers.

bioRxiv : the preprint server for biology·2026
Same author

Context-dependent correlations mislead transcriptomic network inference in bulk and single-cell data.

bioRxiv : the preprint server for biology·2026
Same author

WayFindR: investigating feedback in biological pathways.

NAR genomics and bioinformatics·2026
Same author

Clustering Digestive Tract Tumors Using Transcriptomic and Mutation Data.

Cancers·2026
Same author

Insights into semaglutide cardiovascular research: Mechanisms, trials, and frontiers.

European journal of pharmacology·2026

Related Experiment Videos

Sources of variation in false discovery rate estimation include sample size, correlation, and inherent differences

Jiexin Zhang1, Kevin R Coombes

  • 1Department of Bioinformatics and Computational Biology, The University of Texas MD Anderson Cancer Center, Houston, Texas 77030, USA.

BMC Bioinformatics
|January 17, 2013
PubMed
Summary

This study reveals that sample size, gene correlations, and the proportion of differentially expressed genes (DEGs) significantly impact False Discovery Rate (FDR) estimation accuracy. The proportion of DEGs is a crucial, often overlooked, factor in FDR control for high-throughput data analysis.

Related Experiment Videos

Area of Science:

  • Genomics
  • Statistical genetics
  • Bioinformatics

Background:

  • High-throughput technologies generate vast datasets, necessitating robust statistical methods for analyzing differential gene expression.
  • The multiple testing problem arises when identifying significant genes from numerous simultaneous tests.
  • False Discovery Rate (FDR) control is standard for multiple comparisons, but its accuracy can be compromised by correlated gene expression data.

Purpose of the Study:

  • To evaluate the accuracy of False Discovery Rate (FDR) estimation in the context of high-throughput biological data.
  • To identify and quantify the factors influencing the precision and magnitude of FDR estimates.
  • To highlight the impact of gene correlations and the proportion of differentially expressed genes (DEGs) on FDR control.

Main Methods:

  • Analysis of two real-world datasets with patient subgroup resampling to assess FDR estimation variability.
  • Generation of simulated datasets incorporating block correlation structures and realistic noise using the UMPIRE R package.
  • Estimation of FDR using a beta-uniform mixture (BUM) model to examine variations in results.

Main Results:

  • FDR estimation accuracy is primarily influenced by sample size, gene correlations, and the true proportion of DEGs.
  • Sample size and the proportion of DEGs impact both the size and precision of FDR estimates.
  • Gene correlation structures predominantly affect the variability of estimated FDR parameters.

Conclusions:

  • Key factors affecting FDR estimation have been identified and their impact quantified.
  • The proportion of DEGs is a critical determinant of FDR estimation and warrants careful consideration in statistical analyses.
  • Findings underscore the need for nuanced approaches to FDR control, especially when dealing with correlated genomic data.