Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Introduction to z Scores01:06

Introduction to z Scores

11.2K
A z score (or standardized value) is measured in units of the standard deviation. It tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
11.2K
Introduction to z Scores01:05

Introduction to z Scores

1.4K
A z score (or standardized value) is measured in units of the standard deviation. It indicates how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
1.4K
z Scores and Area Under the Curve01:17

z Scores and Area Under the Curve

19.6K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
19.6K
z Scores and Unusual Values01:07

z Scores and Unusual Values

11.0K
The z score is one of the three measures of relative standing. It describes the location of a value in a dataset relative to the mean. z scores are obtained after the standardization of the values in a dataset. The z score for the mean is 0.
 This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data...
11.0K
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT01:25

Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT

503
Calcium-Scoring CT ScanA calcium-scoring CT scan, also known as coronary artery calcium (CAC) scan, detects calcium deposits in the coronary arteries. This test assesses the risk of coronary artery disease (CAD), which can lead to cardiovascular events such as angina, heart failure, and sudden cardiac arrest.A calcium-scoring CT scan is generally recommended for individuals at intermediate risk of CAD without symptoms. It includes:Men aged 40-75 and women aged 50-75: Especially those with a...
503
Social Scripts02:10

Social Scripts

10.3K
People tend to know what behavior is expected of them in specific, familiar settings. A script is a person’s knowledge about the sequence of events expected in a specific setting (Schank & Abelson, 1977). Essentially, scripts are a particular kind of schema, one containing default values for the features within an event. In the restaurant example, the script's features include the props (e.g., tables, menu, food, and money), the roles to be played (e.g., customer and waiter),...
10.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Stochastic approximation EM for large-scale exploratory IRT factor analysis.

Statistics in medicine·2019
See all related articles

Related Experiment Video

Updated: Feb 7, 2026

Isolation of Fidelity Variants of RNA Viruses and Characterization of Virus Mutation Frequency
18:10

Isolation of Fidelity Variants of RNA Viruses and Characterization of Virus Mutation Frequency

Published on: June 16, 2011

30.1K

IRT Scoring and Test Blueprint Fidelity.

Gregory Camilli1

  • 1Rutgers, The State University of New Jersey, New Brunswick, USA.

Applied Psychological Measurement
|July 24, 2018
PubMed
Summary

This study examines if item response theory (IRT) scoring models align with test blueprint content allocation. Findings suggest that scoring weights in the three-parameter logistic (3PL) model may alter intended proportions for lower-ability examinees.

Keywords:
assessmentguessingitem response theorypsychometric theoryscoring

More Related Videos

A Method for High Fidelity Optogenetic Control of Individual Pyramidal Neurons In vivo
13:44

A Method for High Fidelity Optogenetic Control of Individual Pyramidal Neurons In vivo

Published on: September 2, 2013

19.6K
The Ladder Rung Walking Task: A Scoring System and its Practical Application.
09:38

The Ladder Rung Walking Task: A Scoring System and its Practical Application.

Published on: June 12, 2009

26.8K

Related Experiment Videos

Last Updated: Feb 7, 2026

Isolation of Fidelity Variants of RNA Viruses and Characterization of Virus Mutation Frequency
18:10

Isolation of Fidelity Variants of RNA Viruses and Characterization of Virus Mutation Frequency

Published on: June 16, 2011

30.1K
A Method for High Fidelity Optogenetic Control of Individual Pyramidal Neurons In vivo
13:44

A Method for High Fidelity Optogenetic Control of Individual Pyramidal Neurons In vivo

Published on: September 2, 2013

19.6K
The Ladder Rung Walking Task: A Scoring System and its Practical Application.
09:38

The Ladder Rung Walking Task: A Scoring System and its Practical Application.

Published on: June 12, 2009

26.8K

Area of Science:

  • Educational Measurement
  • Psychometrics
  • Statistical Modeling

Background:

  • Test specifications, or blueprints, guide assessment design and content allocation.
  • Item Response Theory (IRT) provides statistical models for test scoring.
  • Standard IRT models use optimal scoring weights derived from item parameters.

Purpose of the Study:

  • To investigate whether standard IRT scoring models accurately reflect the intended content allocation specified in a test blueprint.
  • To assess if the scoring weights generated by IRT models align with the proportion of items designated for each content area in the blueprint.
  • To specifically examine the implications of the three-parameter logistic (3PL) model, where scoring weights are ability-dependent, on intended content representation.

Main Methods:

  • Analysis of scoring weights derived from standard Item Response Theory (IRT) models, specifically the two-parameter logistic (2PL) and three-parameter logistic (3PL) models.
  • Comparison of IRT-generated optimal scoring weights against a predefined set of intended weights based on the proportion of items in each cell of a test blueprint.
  • Focus on the 3PL model to evaluate how ability-dependent scoring weights might influence the representation of content areas.

Main Results:

  • Standard IRT scoring models employ optimal weights that are contingent upon item parameters.
  • The three-parameter logistic (3PL) model's scoring weights are influenced by examinee ability levels.
  • A potential discrepancy exists where the intended content weights, as defined by the test blueprint, may be implicitly altered for examinees with lower ability in the 3PL model.

Conclusions:

  • The alignment between IRT scoring models and test blueprint content allocation requires careful consideration, particularly with complex models like the 3PL.
  • The ability-dependent nature of scoring weights in the 3PL model raises concerns about maintaining intended content representation across all examinee ability levels.
  • Further research is needed to ensure IRT scoring accurately reflects test design specifications, especially for diverse examinee populations.