Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reliability and Validity01:29

Reliability and Validity

Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
Self-Report Tests of Personality01:22

Self-Report Tests of Personality

Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
Ordinal Level of Measurement00:55

Ordinal Level of Measurement

The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Ratio Level of Measurement00:54

Ratio Level of Measurement

The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated. For...
Nominal Level of Measurement00:56

Nominal Level of Measurement

The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. Not every statistical operation can be used with every set of data. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal scale is...
Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Efficacy of Complexity-Based Target Selection for Treating Morphosyntactic Deficits in Children With Developmental Language Disorder and Children With Down Syndrome: A Single-Case Experimental Design.

American journal of speech-language pathology·2024
Same author

Systematic Review of Variables Related to Instruction in Augmentative and Alternative Communication Implementation: Group and Single-Case Design.

American journal of speech-language pathology·2023
Same author

A Meta-Analysis of Video Modeling Interventions to Enhance Job Skills of Autistic Adolescents and Adults.

Autism in adulthood : challenges and management·2023
Same author

Breathing Exercises, Etc.: A Paper Read before the Los Angeles County Medical Society, April 6, 1888.

Hall's journal of health·2022
Same author

Participant characteristics predicting communication outcomes in AAC implementation for individuals with ASD and IDD: a systematic review and meta-analysis.

Augmentative and alternative communication (Baltimore, Md. : 1985)·2022
Same author

A Meta-analysis of Challenging Behavior Interventions for Students with Developmental Disabilities in Inclusive School Settings.

Journal of autism and developmental disorders·2020

Related Experiment Video

Updated: May 13, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
10:39

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning

Published on: August 29, 2025

Reliability of multi-category rating scales.

Richard I Parker1, Kimberly J Vannest, John L Davis

  • 1Texas A&M University at College Station, USA. rparker@tamu.edu

Journal of School Psychology
|March 14, 2013
PubMed
Summary

Inter-rater reliability for multi-category scales used in education is not well understood. This study found that no single reliability index works best, as performance varies by scale type and data grouping.

Area of Science:

  • Educational Psychology
  • Measurement and Statistics

Background:

  • Multi-category scales are increasingly used for monitoring Individualized Education Program (IEP) goals, classroom rules, and Behavior Improvement Plans (BIPs).
  • Assessing the inter-rater reliability of these scales is crucial but understudied, especially compared to traditional data counting methods.

Purpose of the Study:

  • To examine the performance of nine different reliability indices when applied to multi-category scales.
  • To investigate how scale gradation (number of categories) and data grouping affect reliability index values.

Main Methods:

  • A simulation study was conducted using quasi-continuous data (1-30) transformed into six multi-category scales (2, 3, 5, 7, 10, and 15 points).
  • Nine distinct reliability indices were applied to these scales to evaluate their behavior and performance.

More Related Videos

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)
10:58

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)

Published on: August 28, 2021

Related Experiment Videos

Last Updated: May 13, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
10:39

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning

Published on: August 29, 2025

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)
10:58

Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)

Published on: August 28, 2021

Main Results:

  • Each reliability index demonstrated unique performance characteristics, necessitating individualized interpretation.
  • No single 'best' reliability index was identified; most indices proved to be scale-dependent.
  • Index values changed significantly when categories were collapsed, highlighting the impact of scale structure.

Conclusions:

  • Current reliability indices are not universally applicable to multi-category ordinal scales.
  • New guidelines are required for selecting and interpreting reliability measures for these educational assessment tools.
  • Further research is needed to establish optimal methods for assessing inter-rater reliability with ordinal scales.