Problems with using the normal distribution--and ways to improve quality and efficiency of data analysis

Eckhard Limpert1, Werner A Stahel

  • 1ELI-o-Research, Life Sciences, Zurich, Switzerland.

Plos One
|July 23, 2011
PubMed
Summary

The standard method of describing data variation using the arithmetic mean and standard deviation (SD) is often inadequate for skewed distributions. A log-normal approach using the geometric mean and multiplicative standard deviation offers a more accurate representation, improving data interpretation and potentially reducing sample sizes.

Related Concept Videos

Introduction to Normal Distributions01:29

Introduction to Normal Distributions

Standardized test scores often follow a symmetric distribution that can be modeled with the normal distribution, a fundamental concept in statistics. This distribution is particularly useful for interpreting test performance fairly across populations, as it provides a mathematical framework for understanding variability and central tendency in large datasets.From Histogram to Frequency DistributionRaw test data are often displayed using histograms, where the height of each bar represents the...
Distributions to Estimate Population Parameter01:26

Distributions to Estimate Population Parameter

The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
Data: Types and Distribution01:19

Data: Types and Distribution

In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Normal Distribution01:11

Normal Distribution

The normal, a continuous distribution, is the most important of all the distributions. Its graph is a bell-shaped symmetrical curve, which is observed in almost all disciplines. Some of these include psychology, business, economics, the sciences, nursing, and, of course, mathematics. Some instructors may use the normal distribution to help determine students’ grades. Most IQ scores are normally distributed. Often real-estate prices fit a normal distribution. The normal distribution is extremely...
Student t Distribution01:31

Student t Distribution

The population standard deviation is rarely known in many day-to-day examples of statistics. When the sample sizes are large, it is easy to estimate the population standard deviation using a confidence interval, which provides results close enough to the original value. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
The Student t distribution was developed by William S. Goset (1876–1937) of the...
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...