Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Wald-Wolfowitz Runs Test I01:17

Wald-Wolfowitz Runs Test I

727
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
727
Wald-Wolfowitz Runs Test II01:17

Wald-Wolfowitz Runs Test II

305
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
305
Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

283
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
283
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test01:09

Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test

1.8K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.8K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

679
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
679
The Mantel-Cox Log-Rank Test01:19

The Mantel-Cox Log-Rank Test

503
The Mantel-Cox log-rank test is a widely used statistical method for comparing the survival distributions of two groups. It tests whether a statistically significant difference exists in survival times between the groups without assuming a specific distribution for the survival data, making it a non-parametric test. This flexibility makes the log-rank test particularly valuable in medical research and other fields where the timing of an event, such as death or disease recurrence, is of...
503

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Score-Based Tests With Fixed Effects Person Parameters in Item Response Theory: Detecting Model Misspecification Including Differential Item Functioning.

Applied psychological measurement·2026
Same author

Development and internal validation of the patient satisfaction questionnaire: Perception of Quality in Anaesthesia (PQA)-10.

British journal of anaesthesia·2025
Same author

Major depressive disorder in children and adolescents is associated with reduced hair cortisol and anandamide (AEA): cross-sectional and longitudinal evidence from a large randomized clinical trial.

Translational psychiatry·2025
Same author

Score-based tests for parameter instability in ordinal factor models.

The British journal of mathematical and statistical psychology·2025
Same author

Testing measurement invariance in a conditional likelihood framework by considering multiple covariates simultaneously.

Behavior research methods·2025
Same author

Investigating heterogeneity in IRTree models for multiple response processes with score-based partitioning.

The British journal of mathematical and statistical psychology·2024

Related Experiment Video

Updated: Aug 30, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

3.4K

Power Analysis for the Wald, LR, Score, and Gradient Tests in a Marginal Maximum Likelihood Framework: Applications

Felix Zimmer1, Clemens Draxler2, Rudolf Debelak3

  • 1University of Zurich, Zurich, Switzerland. felix.zimmer@uzh.ch.

Psychometrika
|August 27, 2022
PubMed
Summary

New methods for power analysis and sample size planning in item response theory (IRT) models are introduced. These methods, applicable with marginal maximum likelihood estimation, aid in assessing model fit and differential item functioning for large-scale assessments.

Keywords:
Wald testgradient testitem response theorylikelihood ratiomarginal maximum likelihoodpower analysisscore test

More Related Videos

Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
08:27

Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits

Published on: September 27, 2019

7.0K
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.3K

Related Experiment Videos

Last Updated: Aug 30, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

3.4K
Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
08:27

Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits

Published on: September 27, 2019

7.0K
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.3K

Area of Science:

  • Psychometrics
  • Educational Measurement
  • Statistical Modeling

Background:

  • Item response theory (IRT) models are crucial for analyzing educational assessments.
  • Assessing model fit and detecting differential item functioning (DIF) are key challenges.
  • Existing statistical tests (Wald, likelihood ratio, score, gradient statistics) have limitations in power analysis and sample size planning.

Purpose of the Study:

  • To introduce novel methods for power analysis and sample size planning in IRT.
  • To provide tools for hypothesis testing, including overall model fit and DIF detection.
  • To support the application of these methods in practical, large-scale educational assessments.

Main Methods:

  • Development of analytical methods using asymptotic distributions of test statistics under alternative hypotheses.
  • Implementation of a sampling-based approach for computationally intensive scenarios (e.g., >20 items).
  • Utilizing marginal maximum likelihood (MML) estimation for broad applicability across IRT models.

Main Results:

  • Extensive simulation studies validated the proposed methods in diverse settings (Rasch vs. 2PL, DIF, PCM vs. GPCM).
  • Observed test statistic distributions and power aligned well with analytical predictions in large samples.
  • The methods demonstrate accuracy and practical utility for IRT model evaluation.

Conclusions:

  • The proposed power analysis and sample size planning methods enhance the rigor of IRT applications.
  • These methods are valuable for researchers and practitioners in educational measurement.
  • An openly accessible R package is provided for user-friendly implementation.