Related Experiment Video
Updated: Aug 14, 2026

07:13
Measuring Active and Passive Tameness Separately in Mice
Published on: August 10, 2018
Confidence intervals for an effect size measure based on the Mann-Whitney statistic. Part 1: general issues and
1Department of Epidemiology, Statistics and Public Health, Wales College of Medicine, Cardiff University, Heath Park, Cardiff CF14 4XN, UK. newcombe@cf.ac.uk
Statistics in Medicine
|October 28, 2005
Summary
This study introduces a general effect size measure, theta, for comparing two random variables. It
Area of Science:
- Statistics
- Biostatistics
- Machine Learning
Background:
- Comparing distributions of random variables is crucial in various scientific fields.
- Existing effect size measures may not fully capture the separation between distributions.
Purpose of the Study:
- To propose a general measure of effect size, theta, for quantifying distribution separation.
- To introduce an estimator for theta based on the Mann-Whitney U statistic.
- To explore the impact of data discretization and develop confidence interval methods.
Main Methods:
- Defined theta as Pr[Y > X] + (1/2)Pr[Y = X].
- Estimated theta using U/mn, a generalization of the Mann-Whitney U statistic.
- Investigated discretization effects and developed tail-area-based confidence intervals.
Main Results:
- Theta provides a comprehensive measure of distribution separation.
- The U/mn estimator is equivalent to the area under the receiver operating characteristic curve.
- Developed methods applicable to small samples and extreme outcomes.
Conclusions:
- Theta is a versatile and interpretable effect size measure.
- The proposed estimation and inference methods offer practical solutions for comparing distributions.
- The approach is robust to data discretization and suitable for various sample sizes.
Related Concept Videos
Confidence Intervals
An unbiased point estimate is often insufficient to predict a population estimate, such as population mean or population proportion. In this scenario, a confidence interval is used. A confidence interval is an estimate similar to a sample proportion. However, unlike the point estimate which is a single value, the confidence interval contains a range of values. These values have lower and upper limits, known as confidence limits, and can be designated as L1 and L2, respectively.
A confidence...
A confidence...
Interpretation of Confidence Intervals
A confidence interval is a better estimate of the population than a point estimate, as it uses a range of values from a sample instead of a single value.
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Confidence Interval for Estimating Population Mean
A point estimate of the population mean is obtained from a single sample. Such a point estimate does not represent a population well because it needs to account for variability in the population. Single point estimate can also be biased despite the sample being selected randomly. Thus, a point estimate is often unreliable. A confidence interval is needed to reduce this unreliability.
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
One-Way ANOVA: Equal Sample Sizes
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
Uncertainty: Confidence Intervals
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor 't,' or...

