Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
2.5K
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

3.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.4K
Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

200
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
200
Test for Homogeneity01:23

Test for Homogeneity

2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test01:09

Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test

1.6K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.6K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

517
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
517

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Regional variability in the associations between social and health-related risk factors and memory across Europe.

Scientific reports·2026
Same author

Developing and validating a frailty score based on patient-reported outcome 3 months after stroke: A Riksstroke-based study.

PloS one·2026
Same author

Long-term memory effects of an incremental blood pressure intervention in a mortal cohort.

Biometrics·2026
Same author

The bit scale: A metric score scale for unidimensional item response theory models.

Psychometrika·2025
Same author

A Bayesian semi-parametric approach to causal mediation for longitudinal mediators and time-to-event outcomes with application to a cardiovascular disease cohort study.

Biostatistics (Oxford, England)·2025
Same author

Changing Risks, Changing Outcomes: Cardiovascular Trajectories as a Window Into Dementia Prevention.

Neurology·2025

Related Experiment Video

Updated: Jul 9, 2025

A Tablet-Based Curriculum-Based Measurement Protocol for Kindergarten Writing
15:00

A Tablet-Based Curriculum-Based Measurement Protocol for Kindergarten Writing

Published on: February 7, 2025

587

Efficiency Analysis of Item Response Theory Kernel Equating for Mixed-Format Tests.

Joakim Wallmark1, Maria Josefsson1, Marie Wiberg1

  • 1Department of Statistics, USBE, Umeå University, Sweden.

Applied Psychological Measurement
|November 29, 2023
PubMed
Summary

Item Response Theory (IRT) kernel equating performs well for mixed-format tests, showing minimal differences compared to IRT observed score equating. IRT presmoothing offers smaller standard errors, making it a promising method for test equating.

Keywords:
item response theorykernel equatinglog-linear modelspresmoothingsimulation

More Related Videos

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.2K
Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

5.8K

Related Experiment Videos

Last Updated: Jul 9, 2025

A Tablet-Based Curriculum-Based Measurement Protocol for Kindergarten Writing
15:00

A Tablet-Based Curriculum-Based Measurement Protocol for Kindergarten Writing

Published on: February 7, 2025

587
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.2K
Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

5.8K

Area of Science:

  • Psychometrics
  • Educational Measurement
  • Statistical Modeling

Background:

  • Equating mixed-format tests presents unique challenges.
  • Item Response Theory (IRT) kernel equating is a potential solution.
  • Comparison with existing methods like IRT observed score equating and log-linear kernel equating is necessary.

Purpose of the Study:

  • To evaluate the performance of IRT kernel equating for mixed-format tests.
  • To compare IRT kernel equating against IRT observed score equating and log-linear kernel equating.
  • To assess the impact of IRT models versus non-IRT models in data simulation on equating accuracy.

Main Methods:

  • Simulations and real data applications were used for comparison.
  • Equivalent Groups (EG) and Non-Equivalent Groups with Anchor Test (NEAT) designs were employed.
  • Data were simulated both with and without IRT models to test robustness.

Main Results:

  • IRT kernel equating and IRT observed score equating showed minimal differences in equated scores and standard errors.
  • IRT presmoothing resulted in smaller standard errors of equating compared to log-linear presmoothing.
  • IRT-based methods were less biased when data followed IRT models, while log-linear equating showed less bias for non-IRT data.

Conclusions:

  • IRT kernel equating demonstrates significant promise for equating mixed-format tests.
  • The choice of presmoothing method (IRT vs. log-linear) impacts standard error.
  • IRT-based equating methods are preferable when test data align with IRT assumptions.