Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reliability and Validity01:29

Reliability and Validity

13.0K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.0K
Comparing Experimental Results: Student's t-Test01:09

Comparing Experimental Results: Student's t-Test

1.8K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
1.8K
Kendall's Coefficient of Concordance01:20

Kendall's Coefficient of Concordance

498
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
498
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

276
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
276
Statistical Analysis: Overview01:11

Statistical Analysis: Overview

7.2K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
7.2K
Behrens–Fisher Test00:57

Behrens–Fisher Test

126
The Behrens-Fisher test is a statistical method designed to address the Behrens-Fisher problem, which arises when comparing the means of two normally distributed populations with unequal variances. Unlike the Student's t-test, which assumes equal variances, the Behrens-Fisher test allows for mean comparison without this restrictive assumption. This flexibility makes it particularly valuable in scenarios where two independent samples exhibit normality but lack variance homogeneity.
This test...
126

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The Reward of Virtue: Examining Roles of Patient Gratitude, Physician Affect, and Physician Rumination in Linking Patient-Centered Communication and Physician Turnover.

Health communication·2026
Same author

Discussing diseases in everyday talk: Examining the roles of medical and phatic patient-provider communication in promoting Chinese patients' healthy lifestyle behaviors.

Patient education and counseling·2026
Same author

Willing or reluctant to share health data? A moderated mediation analysis of wearable device usage and data-sharing intentions among older adults.

Digital health·2026
Same author

How do health information seeking and scanning motivate older adults to use patient portals? A mediation analysis based on the technology acceptance model.

Journal of health psychology·2026
Same author

Multifunctional Online Medical Record Use and Patient Empowerment: Examining the Mediating Role of Patient-Centered Communication Across the Life Span in the Greater China Region.

Health communication·2026
Same author

Factors associated with healthy lifestyles in Chinese older adults with chronic conditions: A comparison of online and offline social support, mediated by health management self-efficacy and moderated by online patient-centered communication.

Digital health·2026

Related Experiment Video

Updated: Aug 30, 2025

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
08:40

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity

Published on: June 12, 2019

7.5K

Interrater reliability estimators tested against true interrater reliabilities.

Xinshu Zhao1, Guangchao Charles Feng2, Song Harris Ao2

  • 1Department of Communication, Faculty of Social Sciences, University of Macau, Taipa, Macao. xszhao@um.edu.mo.

BMC Medical Research Methodology
|August 29, 2022
PubMed
Summary

Simple percent agreement, despite criticism, accurately predicts interrater reliability. Popular chance-adjusted indices like Cohen's κ and Krippendorff's α underperform, suggesting a need for new reliability metrics.

Keywords:
Cohen’s kappaIntercoder reliabilityInterrater reliabilityKrippendorff’s alphaReconstructed experiment

More Related Videos

A Protocol of Manual Tests to Measure Sensation and Pain in Humans
07:28

A Protocol of Manual Tests to Measure Sensation and Pain in Humans

Published on: December 19, 2016

21.1K
Assessment of Child Anthropometry in a Large Epidemiologic Study
09:36

Assessment of Child Anthropometry in a Large Epidemiologic Study

Published on: February 2, 2017

27.2K

Related Experiment Videos

Last Updated: Aug 30, 2025

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
08:40

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity

Published on: June 12, 2019

7.5K
A Protocol of Manual Tests to Measure Sensation and Pain in Humans
07:28

A Protocol of Manual Tests to Measure Sensation and Pain in Humans

Published on: December 19, 2016

21.1K
Assessment of Child Anthropometry in a Large Epidemiologic Study
09:36

Assessment of Child Anthropometry in a Large Epidemiologic Study

Published on: February 2, 2017

27.2K

Area of Science:

  • Statistics
  • Psychometrics
  • Medical Research

Background:

  • Interrater reliability measures agreement between raters, crucial for data quality in various fields.
  • Existing indices for interrater reliability lack consensus, with debates on their appropriateness.
  • Percent agreement (ao) is simple but flawed for not accounting for chance agreement.

Purpose of the Study:

  • To experimentally evaluate the performance of seven prominent interrater reliability indices.
  • To assess the impact of rating category, distribution skew, and task difficulty on these indices.
  • To compare the accuracy of established and newer indices in predicting and approximating reliability.

Main Methods:

  • A controlled experiment with 384 rating sessions was conducted.
  • Seven well-known interrater reliability indices were tested.
  • The study analyzed the influence of category, skew, and difficulty on index performance.

Main Results:

  • Percent agreement (ao) was the most accurate predictor of reliability (r2=.84).
  • Scott's π, Cohen's κ, and Krippendorff's α indices underperformed significantly.
  • Gwet's AC1 index showed strong performance as a predictor and approximator.

Conclusions:

  • Current chance-adjusted indices may be flawed due to assumptions about rater behavior.
  • Indices should potentially rely on task difficulty rather than category or skew.
  • Further empirical studies are needed to validate findings and develop improved reliability indices.