Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Understanding Deception01:14

Understanding Deception

181
Deception is a pervasive aspect of human communication. Empirical studies have shown that most individuals engage in some form of deceit on a daily basis, with approximately 20% of social exchanges involving deceptive elements. Lying follows a developmental trajectory, peaking during adolescence and declining with age, possibly due to the maturation of cognitive control and social accountability.Cognitive and Social Factors in Deception DetectionDespite its prevalence, accurately detecting...
181
Calibration Curves: Correlation Coefficient01:10

Calibration Curves: Correlation Coefficient

5.0K
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the...
5.0K
Kendall's Coefficient of Concordance01:20

Kendall's Coefficient of Concordance

1.1K
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
1.1K
Theory of Attribution I: Correspondent Inference Theory01:15

Theory of Attribution I: Correspondent Inference Theory

627
Correspondent inference theory, proposed by Jones and Davis in 1965, seeks to explain how individuals infer stable personality traits from observed behaviors. It suggests that people attribute actions to underlying dispositions rather than external circumstances, particularly when the behavior appears intentional and socially significant.Voluntary Behavior and Dispositional AttributionAccording to this theory, individuals are more likely to attribute behavior to personal traits when it appears...
627
Reliability and Validity01:29

Reliability and Validity

14.2K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
14.2K
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

16.7K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
16.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Tool for Agreement and Alignment Analysis in Binary Rating Tasks: The R Package scindex.

Applied psychological measurement·2026
Same author

babebi: An R Package for Bayesian Estimation and Validation in Small-N Two-Rater Pre-Post Designs.

Applied psychological measurement·2026
Same author

Agreement and Alignment in Binary Rating Tasks: Strategic Convergence as an Equilibrium Outcome.

Educational and psychological measurement·2026
See all related articles

Related Experiment Video

Updated: Feb 20, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

3.0K

From Agreement to Epistemic Alignment: A Signal Detection-Theoretic Model of Inter-Rater Reliability.

Irene Gianeselli1

  • 1Free University of Bozen-Bolzano, Italy.

Educational and Psychological Measurement
|February 19, 2026
PubMed
Summary

Cohen

Keywords:
Cohen’s Kappainter-rater reliabilitypsychometricssignal detection theorysimulation study

More Related Videos

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
08:40

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity

Published on: June 12, 2019

7.9K

Related Experiment Videos

Last Updated: Feb 20, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
08:12

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments

Published on: March 1, 2022

3.0K
Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
08:40

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity

Published on: June 12, 2019

7.9K

Area of Science:

  • Psychology
  • Statistics
  • Cognitive Science

Background:

  • Inter-rater reliability is often measured using Cohen's kappa (κ), which quantifies agreement but doesn't model judgment processes.
  • Cohen's kappa is sensitive to prevalence, difficulty, and differing rater criteria, leading to misinterpretations of diagnostic accuracy.

Purpose of the Study:

  • To reframe inter-rater reliability using signal detection theory (SDT).
  • To introduce the Strategic Convergence Index (SCI) as a measure of rater decision threshold alignment within an SDT framework.

Main Methods:

  • Developed a signal detection-theoretic generative model for categorical judgments.
  • Introduced the Strategic Convergence Index (SCI) to quantify convergence in rater decision thresholds.
  • Conducted Monte Carlo simulations to compare Cohen's kappa and SCI under varying conditions.

Main Results:

  • Cohen's kappa varies with prevalence and discriminability, even with constant decision policies.
  • The Strategic Convergence Index (SCI) selectively measures epistemic alignment and is invariant to these factors.
  • SCI can be recovered as a stable property even with uncertainty about the true state.

Conclusions:

  • Cohen's kappa is a bounded transformation of strategic variance, not a direct measure of epistemic alignment.
  • The Strategic Convergence framework decomposes reliability into outcome agreement and process alignment.
  • This framework advances reliability assessment from description toward explanation by linking statistics to a generative model.