Related Experiment Video
Updated: Aug 2, 2026

07:41
Modeling the Size Spectrum for Macroinvertebrates and Fishes in Stream Ecosystems
Published on: July 30, 2019
Sample size-based indication of normality in lognormally distributed populations
Applied Occupational and Environmental Hygiene
|August 3, 1999
Summary
Distinguishing between Normal and Lognormal distributions in occupational hygiene sampling requires more samples as the geometric standard deviation (GSD) decreases. Insufficient samples hinder accurate extreme event estimations.
Area of Science:
- Environmental Health
- Occupational Hygiene
- Statistical Modeling
Background:
- Occupational and environmental hygiene sampling often yields small sample sizes.
- Estimates from small samples are used for extreme event predictions, influenced by exposure distributions.
Purpose of the Study:
- To investigate sample size limitations in distinguishing between Normal and Lognormal distributions.
- To assess the impact of geometric standard deviation (GSD) on distribution identification.
Main Methods:
- Generated synthetic samples (5-250) from Lognormal distributions with varying GSDs (1.25-7.00).
- Applied Shapiro and Wilk's W Test to assess goodness of fit for Normality.
- Repeated simulations to ensure robustness of results.
Main Results:
- The number of samples needed to differentiate distributions inversely correlated with GSD.
- 169 samples were required for 90% distinction at GSD=1.25 (alpha=0.05).
- Fewer samples (25 and 15) were needed for GSDs of 2.00 and 4.00, respectively.
Conclusions:
- Small sample sizes pose significant challenges in distinguishing between Normal and Lognormal distributions in hygiene sampling.
- This difficulty can lead to inaccurate estimations of extreme exposure events.
- Statistical power is crucial for reliable exposure assessment in occupational and environmental health.
Related Concept Videos
Central Limit Theorem
The central limit theorem, abbreviated as clt, is one of the most powerful and useful ideas in all of statistics. The central limit theorem for sample means says that if you repeatedly draw samples of a given size and calculate their means, and create a histogram of those means, then the resulting histogram will tend to have an approximate normal bell shape. In other words, as sample sizes increase, the distribution of means follows the normal distribution more closely.
The sample size, n, that...
The sample size, n, that...
Estimating Population Mean with Known Standard Deviation
To construct a confidence interval for a single unknown population mean μ, where the population standard deviation is known, we need sample mean as an estimate for μ and we need the margin of error. Here, the margin of error (EBM) is called the error bound for a population mean (abbreviated EBM). The sample mean is the point estimate of the unknown population mean μ.
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate + error bound)
The...
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate + error bound)
The...
Distributions to Estimate Population Parameter
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
Student t Distribution
The population standard deviation is rarely known in many day-to-day examples of statistics. When the sample sizes are large, it is easy to estimate the population standard deviation using a confidence interval, which provides results close enough to the original value. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
The Student t distribution was developed by William S. Goset (1876–1937) of the...
The Student t distribution was developed by William S. Goset (1876–1937) of the...
The Anderson-Darling Test
The Anderson-Darling test is a statistical method used to determine whether a data sample is likely drawn from a specific theoretical distribution. Unlike parametric tests, it does not require assumptions about specific parameters of the distribution. Instead, it compares the sample's empirical cumulative distribution function (ECDF) with the cumulative distribution function (CDF) of the hypothesized distribution. Critical values for the test are specific to the chosen distribution rather than...
Testing a Claim about Mean: Unknown Population SD
A complete procedure of testing a hypothesis about a population mean when the population standard deviation is unknown is explained here.
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used; instead...
Estimating a population mean requires the samples to be approximately normally distributed. The data should be collected from the randomly selected samples having no sampling bias. There is no specific requirement for sample size. But if the sample size is less than 30, and we don't know the population standard deviation, a different approach is used; instead...

