Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reliability and Validity01:29

Reliability and Validity

Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
Blind Procedures02:07

Blind Procedures

Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which child was...
Variability: Analysis01:11

Variability: Analysis

Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The Moderation Effect of Basic Need Satisfaction in the Relationship Between Drinking Motives and Drinking Outcomes.

Substance use & misuse·2026
Same author

Clinical updates: Assessment and management of suicidal ideation in adults.

BMJ (Clinical research ed.)·2026
Same author

Drinking to belong: how loneliness fuels alcohol-related consequences.

The Journal of social psychology·2026
Same author

A randomized controlled trial of injunctive norms feedback with and without motivational interviewing to reduce alcohol use and negative consequences among college students.

Journal of consulting and clinical psychology·2026
Same author

Expanding perspectives on reducing harms from drinking in college: A qualitative study.

Psychology of addictive behaviors : journal of the Society of Psychologists in Addictive Behaviors·2026
Same author

Mechanisms of change for two brief alcohol interventions: Testing theoretical mediators for counter attitudinal advocacy and personalized feedback intervention effects.

Alcohol, clinical & experimental research·2025

Related Experiment Video

Updated: Jul 9, 2026

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
08:40

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity

Published on: June 12, 2019

Differences in inter-rater reliability and accuracy for a treatment adherence scale.

Salene M Wu1, Ursula Whiteside, Clayton Neighbors

  • 1University of Washington, Seattle, Washington 98195, WUSA. smw9@u.washington.edu

Cognitive Behaviour Therapy
|December 1, 2007
PubMed
Summary

Inter-rater reliability and accuracy are distinct measures of rater performance. This study found that while accuracy and reliability can be influenced by factors like therapist behavior, neither can be assumed to exceed the other.

Related Experiment Videos

Last Updated: Jul 9, 2026

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
08:40

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity

Published on: June 12, 2019

Area of Science:

  • Psychology
  • Clinical Psychology
  • Research Methodology

Background:

  • Inter-rater reliability and accuracy are key metrics for evaluating rater performance.
  • Inter-rater reliability is often used as a proxy for accuracy, despite documented conceptual and empirical differences.
  • Understanding the distinction and relationship between these measures is crucial for reliable research and clinical assessment.

Purpose of the Study:

  • To compare inter-rater reliability and accuracy in assessing therapist adherence to cognitive behavioral therapy (CBT).
  • To identify factors influencing the reliability and accuracy of these ratings.
  • To evaluate the effectiveness of consensus and composite ratings compared to individual ratings.

Main Methods:

  • Paired undergraduate raters assessed therapist adherence from videotaped CBT sessions.
  • Ratings were compared against expert-generated criterion ratings and between raters using intraclass correlation.
  • Therapist-specific factors and behavioral frequencies/intensities were analyzed for their impact on ratings.

Main Results:

  • Inter-rater reliability was marginally, though not significantly, higher than accuracy (p = 0.09).
  • The specific therapist and the frequency/intensity of their behaviors significantly impacted both reliability and accuracy.
  • Consensus ratings were more accurate than individual ratings, but composite ratings did not offer further accuracy improvement.

Conclusions:

  • Accuracy and inter-rater reliability are not interchangeable and are influenced by multiple factors, including the subject being rated.
  • Therapist characteristics and the nature of observable behaviors significantly affect rater performance.
  • Composite ratings may enhance accuracy over individual ratings, while consensus ratings may not justify the additional time investment.