Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Bootstrapping01:24

Bootstrapping

691
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
691
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test01:09

Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test

3.0K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
3.0K
Friedman Two-way Analysis of Variance by Ranks01:21

Friedman Two-way Analysis of Variance by Ranks

333
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
333
Response Surface Methodology01:16

Response Surface Methodology

325
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
325
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

5.6K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
5.6K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

805
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
805

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Examining Differential Rater Functioning and Bias in the Holistic Review of Residency Applications.

Journal of graduate medical education·2026
Same author

Mental Rotation Performance: Contribution of Item Features to Difficulties and Functional Adaptation.

Journal of Intelligence·2025
Same author

Exploring the Impact of Missing Data on Residual-Based Dimensionality Analysis for Measurement Models.

Educational and psychological measurement·2023
Same author

Comparing Person-Fit and Traditional Indices Across Careless Response Patterns in Surveys.

Applied psychological measurement·2023
Same author

Does Sparseness Matter? Examining the Use of Generalizability Theory and Many-Facet Rasch Measurement in Sparse Rating Designs.

Applied psychological measurement·2023
Same author

Detecting Rating Scale Malfunctioning With the Partial Credit Model and Generalized Partial Credit Model.

Educational and psychological measurement·2023

Related Experiment Video

Updated: Oct 19, 2025

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K

An Iterative Parametric Bootstrap Approach to Evaluating Rater Fit.

Wenjing Guo1, Stefanie A Wind1

  • 1The University of Alabama, Tuscaloosa, USA.

Applied Psychological Measurement
|September 27, 2021
PubMed
Summary

Researchers often use rule-of-thumb values for rater fit statistics, but these can be inaccurate. An iterative bootstrap procedure offers a more reliable method for detecting problematic rater patterns in performance assessments.

Keywords:
false-positive ratesparametric bootstrap methodrater-mediated assessmenttrue-positive rate

More Related Videos

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.7K
Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
08:27

Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits

Published on: September 27, 2019

7.0K

Related Experiment Videos

Last Updated: Oct 19, 2025

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.2K
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

2.7K
Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
08:27

Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits

Published on: September 27, 2019

7.0K

Area of Science:

  • Educational measurement
  • Psychometrics
  • Item Response Theory (IRT)

Background:

  • Performance assessments utilize measurement theory models to identify rater discrepancies.
  • Infit and outfit mean square error (MSE) statistics are commonly used to detect problematic scoring patterns.
  • Traditional rule-of-thumb critical values for MSE statistics may not be appropriate in many empirical settings.

Purpose of the Study:

  • To evaluate the effectiveness of parametric bootstrapped critical values for detecting rater misfit.
  • To assess the false-positive and true-positive rates of infit and outfit MSE statistics using a simulation study.
  • To propose and evaluate an iterative parametric bootstrap procedure to improve rater misfit detection.

Main Methods:

  • A simulation study was conducted to assess the performance of traditional parametric bootstrap and rule-of-thumb critical values.
  • An iterative parametric bootstrap procedure was developed and implemented.
  • False-positive and true-positive rates of infit and outfit MSE statistics were analyzed.

Main Results:

  • Traditional parametric bootstrap and rule-of-thumb methods exhibited inflated false-positive rates and low true-positive rates for rater misfit detection.
  • The proposed iterative parametric bootstrap procedure demonstrated better control over false-positive rates.
  • The iterative procedure also yielded higher true-positive rates compared to traditional methods.

Conclusions:

  • Standard methods for interpreting infit and outfit MSE statistics in performance assessments have limitations.
  • An iterative parametric bootstrap procedure provides a more accurate and reliable approach for identifying rater misfit.
  • This improved method enhances the validity of performance assessment evaluations.