Bias in the estimation of false discovery rate in microarray studies

Yudi Pawitan1, Karuturi R Krishna Murthy, Stefan Michiels

  • 1Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, Sweden. yudi.pawitan@meb.ki.se

Abstract

Insights

This study introduces a new mixture model to accurately estimate the proportion of non-differentially expressed genes (pi(0)) in microarray analysis. This improved estimation reduces bias and increases statistical power in identifying significant gene expression changes.

Area of Science:

  • Genomics
  • Statistical Bioinformatics

Background:

  • The false discovery rate (FDR) is crucial for microarray studies, but its accuracy depends on estimating pi(0), the proportion of non-differentially expressed genes.
  • Current methods often overestimate pi(0) and FDR, leading to reduced statistical power in identifying significant genes.

Purpose of the Study:

  • To develop an improved method for estimating pi(0) in two-sample microarray comparisons.
  • To address the bias in standard pi(0) estimation, particularly when pi(0) is far from 1 or test statistic non-centrality parameters are small.

Main Methods:

  • Derived a natural mixture model for the test statistic in two-sample comparisons.
  • Developed an explicit bias formula for standard pi(0) estimation.
  • Proposed a practical, likelihood-based procedure for improved pi(0) estimation using the mixture model.

Main Results:

  • Identified significant bias in standard pi(0) estimation under specific conditions.
  • Demonstrated that mixture-model estimates of pi(0) are less biased than standard estimates through simulations.
  • Explained discrepancies between non-parametric and model-based pi(0) estimates.

Conclusions:

  • The proposed mixture model and likelihood-based procedure offer a more accurate estimation of pi(0) for microarray data.
  • This improved estimation enhances the reliability of FDR calculations and increases statistical power.
  • An R-package, OCplus, is available for implementing these methods.

Related Concept Videos

DNA Microarrays02:34

DNA Microarrays

Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
Bias01:22

Bias

Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...