Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Introduction to z Scores01:05

Introduction to z Scores

964
A z score (or standardized value) is measured in units of the standard deviation. It indicates how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
964
Introduction to z Scores01:06

Introduction to z Scores

10.8K
A z score (or standardized value) is measured in units of the standard deviation. It tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
10.8K
Introduction to the Sign Test01:10

Introduction to the Sign Test

1.2K
The sign test is an important tool in nonparametric statistics, offering a straightforward yet effective method for analyzing matched pairs, nominal data, or hypotheses concerning the median of a population. It transforms data points into positive or negative signs, avoiding the need for assumptions about data distribution and instead focusing on the direction of change. It is particularly valuable when data does not conform to the normal distribution requirements of many parametric tests. For...
1.2K
Comparing Experimental Results: Student's t-Test01:09

Comparing Experimental Results: Student's t-Test

4.5K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.5K
Wald-Wolfowitz Runs Test II01:17

Wald-Wolfowitz Runs Test II

460
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
460
z Scores and Unusual Values01:07

z Scores and Unusual Values

10.8K
The z score is one of the three measures of relative standing. It describes the location of a value in a dataset relative to the mean. z scores are obtained after the standardization of the values in a dataset. The z score for the mean is 0.
 This score indicates how far a value is from the mean in terms of standard deviation. For example, if a data value has a z score of +1, the researcher can infer that the particular data value is one standard deviation above the mean. If another data...
10.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Conditional Reliability of Weighted Test Scores on a Bounded <i>D</i>-Scale.

Educational and psychological measurement·2025
Same author

The Dominant Trait Profile Method of Scoring Multidimensional Forced-Choice Questionnaires.

Educational and psychological measurement·2025
Same author

Scalable Career & Educational Growth Transitions: Developing a Mentoring Network & Engaging with an Online Mentoring Platform.

The chronicle of mentoring & coaching·2024
Same author

Latent <i>D</i>-Scoring Modeling: Estimation of Item and Person Parameters.

Educational and psychological measurement·2023
Same author

The Response Vector for Mastery Method of Standard Setting.

Educational and psychological measurement·2022
Same author

Testing for Differential Item Functioning Under the <i>D</i>-Scoring Method.

Educational and psychological measurement·2022

Related Experiment Video

Updated: Dec 15, 2025

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

1.1K

The Delta-Scoring Method of Tests With Binary Items: A Note on True Score Estimation and Equating.

Dimiter M Dimitrov1,2

  • 1George Mason University, Fairfax, VA, USA.

Educational and Psychological Measurement
|July 14, 2020
PubMed
Summary

This study introduces delta scoring (D-scoring) for binary test items, offering new methods for scaling and equating without needing item response theory calibration. This approach is being piloted for large-scale assessments in Saudi Arabia.

Keywords:
assessmentdelta-scoringtest equatingtest scoringtesting

More Related Videos

Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

6.1K
Measurement & Analysis of the Temporal Discrimination Threshold Applied to Cervical Dystonia
10:05

Measurement & Analysis of the Temporal Discrimination Threshold Applied to Cervical Dystonia

Published on: January 27, 2018

10.1K

Related Experiment Videos

Last Updated: Dec 15, 2025

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

1.1K
Computerized Adaptive Testing System of Functional Assessment of Stroke
05:21

Computerized Adaptive Testing System of Functional Assessment of Stroke

Published on: January 7, 2019

6.1K
Measurement & Analysis of the Temporal Discrimination Threshold Applied to Cervical Dystonia
10:05

Measurement & Analysis of the Temporal Discrimination Threshold Applied to Cervical Dystonia

Published on: January 27, 2018

10.1K

Area of Science:

  • Educational Measurement
  • Psychometrics
  • Statistical Modeling

Background:

  • Traditional test scoring and equating methods can be complex.
  • Item Response Theory (IRT) calibration is often required for advanced scoring techniques.
  • Delta scoring (D-scoring) offers an alternative approach for binary items.

Purpose of the Study:

  • To present new developments in delta scoring (D-scoring) methodology for binary test items.
  • To introduce procedures for scaling, equating, and estimating D scores, true values, and standard errors.
  • To present a D-scoring approach that does not require Item Response Theory (IRT) calibration.

Main Methods:

  • Development of new procedures for scaling and equating binary test items using D-scoring.
  • Formulation of the item response function within the D-scoring framework.
  • Estimation of true values and standard errors for D scores.

Main Results:

  • The proposed D-scoring methodology provides procedures for scaling and equating tests with binary items.
  • The new approach successfully estimates true values and standard errors of D scores.
  • The methodology avoids the need for IRT calibration, simplifying the process.

Conclusions:

  • The enhanced D-scoring methodology offers a viable alternative for scoring and equating binary tests.
  • This approach is suitable for large-scale assessments and is currently under piloting.
  • The elimination of IRT calibration makes D-scoring more accessible.