Related Experiment Video
Updated: May 17, 2026

10:10
Three Differential Expression Analysis Methods for RNA Sequencing: limma, EdgeR, DESeq2
Published on: September 18, 2021
The distribution-based p-value for the outlier sum in differential gene expression analysis
Lin-An Chen1, Dung-Tsa Chen, Wenyaw Chan
1Institute of Statistics, National Chiao Tung University , Hsinchu , Taiwan lachen@stat.nctu.edu.tw.
Biometrika
|October 11, 2012
Summary
A new outlier sum method efficiently detects outlier genes in disease samples. This statistical approach offers improved power for identifying genes with unusual expression patterns in DNA microarray data.
Area of Science:
- Genomics
- Statistical Genetics
- Bioinformatics
Background:
- Outlier sums were previously proposed for outlier gene detection.
- Existing methods lacked formal statistical inference and distributional properties.
Purpose of the Study:
- To propose a novel outlier sum for enhanced outlier gene detection.
- To develop the asymptotic distribution theory and formulate a p-value for the new method.
- To assess the efficiency of the proposed method compared to existing techniques.
Main Methods:
- Development of a new outlier sum statistic.
- Derivation of asymptotic distribution theory and analytic p-value formulation.
- Power comparison with existing outlier sum methods using large-sample theory.
Main Results:
- The proposed outlier sum method demonstrates superior efficiency in detecting outlier genes.
- The method was successfully applied to DNA microarray data from breast tumor samples.
- Formal statistical inference and distributional properties were established for the new outlier sum.
Conclusions:
- The novel outlier sum provides a statistically robust and efficient approach for outlier gene detection.
- This method enhances the analysis of genomic data, particularly in identifying disease-related gene expression anomalies.
- The findings suggest improved diagnostic and research capabilities in genomics.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
P-value
P-value is one of the most crucial concepts in statistics.
P-value stands for the probability value. P-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme as the results obtained from the given sample.
A large P-value calculated from the data indicates to not reject the null hypothesis. But a higher P-value does not mean that the null hypothesis is true. The smaller the P-value, the more unlikely...
P-value stands for the probability value. P-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme as the results obtained from the given sample.
A large P-value calculated from the data indicates to not reject the null hypothesis. But a higher P-value does not mean that the null hypothesis is true. The smaller the P-value, the more unlikely...
Detection of Gross Error: The Q Test
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
What Are Outliers?
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
Outliers and Influential Points
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the vertical...
Significance Testing: Overview
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
