Sample size, statistical power, and false conclusions in infant looking-time research

Lisa M Oakes1

  • 1UC Davis.

Insights

Infant research often uses small sample sizes, leading to unreliable results. This study shows that small samples in infant looking time studies increase false positives and negatives, impacting research conclusions.

Area of Science:

  • Developmental Psychology
  • Cognitive Science
  • Infant Research Methodologies

Background:

  • Infant research, particularly using looking time measures, is characterized by small sample sizes due to recruitment and logistical challenges.
  • Many published infant looking time studies utilize fewer than 24 infants per condition, raising concerns about statistical power.

Purpose of the Study:

  • To investigate the impact of small sample sizes on statistical power and research outcomes in infant looking time studies.
  • To evaluate the reliability of conclusions drawn from studies with limited participant numbers.

Main Methods:

  • Analysis of three large infant datasets (>30 infants per condition).
  • Simulation of 1000 subsamples with sizes ranging from 8 to 24 infants.
  • Systematic examination of how varying sample sizes affect statistical results and conclusion validity.

Main Results:

  • Studies with small sample sizes (8-24 infants) demonstrated high variability in outcomes, even when original large samples yielded clear results.
  • Simulated small samples frequently produced both false positive and false negative findings.
  • Low statistical power in typical infant looking time studies contributes to unreliable research findings.

Conclusions:

  • Current small sample sizes in infant looking time research compromise study power and lead to erroneous conclusions.
  • The variability introduced by small samples necessitates a re-evaluation of published findings and research practices.
  • Exploring emerging solutions is crucial to enhance the reliability and validity of infant research.

Related Concept Videos

Statistical Significance01:50

Statistical Significance

Once data is collected from both the experimental and the control groups, a statistical analysis is conducted to find out if there are meaningful differences between the two groups. A statistical analysis determines how likely any difference found is due to chance (and thus not meaningful). In psychology, group differences are considered meaningful, or significant, if the odds that these differences occurred by chance alone are 5 percent or less. Stated another way, if we repeated this...
22.5K
Errors In Hypothesis Tests01:14

Errors In Hypothesis Tests

When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
6.1K
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
620
Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.2K