Related Experiment Video
Updated: May 29, 2026

Qualitative and Quantitative Validation of Tools with Rating Scales Aimed at Assessing the Quality of University Service-Learning
Published on: August 29, 2025
Rating scales and Rasch measurement
1Graduate School of Education, The University of Western Australia, M428, 35 Stirling Highway, Crawley, Western Australia, 6009, Australia. david.andrich@uwa.edu.au
This article examines two competing ways to analyze rating scales in research. It highlights how different statistical approaches lead to distinct conclusions about how we measure human traits and behaviors.
Area of Science:
- Psychometrics and Rasch measurement research
- Statistical modeling in social sciences
Background:
Researchers often lack direct physical tools to quantify abstract properties like intensity or quality. This gap motivated the widespread adoption of rating scales across diverse scientific disciplines. Prior work has established that these scales rely on ordered categories to capture variations in degree. Classical test theory traditionally assigns simple integer scores to these categories. However, this elementary approach often fails to account for the probabilistic nature of human responses. Modern test theory introduced sophisticated models to address these limitations through person and category parameters. That uncertainty drove the development of two distinct, yet conflicting, analytical frameworks. No prior work had resolved the fundamental incompatibility between these competing paradigms until now.
Purpose Of The Study:
The aim of this article is to clarify the incompatible differences between two major paradigms for analyzing rating scales. Researchers frequently struggle with selecting the appropriate model for their specific measurement needs. This study addresses the confusion surrounding the use of statistical modeling versus experimental measurement. The authors seek to explain how these frameworks influence the design of assessment instruments. By focusing on these two paradigms, the article provides a clear comparison of their underlying logic. This work aims to guide substantive researchers and psychometricians in making informed decisions. The authors intend to show that these paradigms are not interchangeable in practice. Ultimately, the study provides a foundation for better understanding the implications of model choice on research inferences.
Main Methods:
The authors employ a comparative review approach to contrast two dominant psychometric frameworks. They evaluate the theoretical foundations of statistical modeling against those of experimental measurement. This review focuses on identifying incompatible points between these two methodologies. The authors avoid exhaustive lists of all available models to maintain clarity. Instead, they examine the core logic governing how each paradigm treats rating categories. A illustrative example serves to demonstrate the practical divergence in analytical outcomes. The investigation emphasizes the implications for researcher roles in instrument design. This systematic comparison provides a framework for understanding how model selection impacts scientific inferences.
Main Results:
The analysis reveals that the statistical modeling and experimental measurement paradigms remain fundamentally incompatible. These frameworks diverge significantly in their treatment of person and category parameters. The authors demonstrate that these differences lead to distinct roles for researchers during instrument development. Statistical modeling often prioritizes data fit, while experimental measurement emphasizes the construction of invariant scales. The study shows that these choices directly influence the validity of inferences drawn from rating data. The authors provide an example to highlight how these paradigms interpret the same responses differently. This comparison clarifies why researchers cannot easily switch between these two analytical approaches. The findings indicate that the selection of a paradigm is a critical decision in psychometric research.
Conclusions:
The authors argue that choosing between statistical modeling and experimental measurement paradigms alters research outcomes. These frameworks demand different levels of involvement from both substantive experts and psychometricians. The analysis demonstrates that these paradigms remain incompatible on several key points. Researchers must recognize that model selection dictates the validity of their subsequent inferences. The study clarifies how these choices influence the design of future assessment instruments. By highlighting these differences, the authors provide a guide for navigating complex measurement challenges. The findings suggest that the role of the investigator changes depending on the chosen paradigm. Ultimately, this work emphasizes the necessity of aligning analytical methods with the underlying goals of the measurement process.
Frequently Asked Questions
The researchers propose that the statistical modeling paradigm focuses on fitting data to predefined distributions, whereas the experimental measurement paradigm prioritizes the construction of invariant scales. These approaches differ in how they handle person and category parameters during the estimation process.
The authors discuss rating scales, which utilize ordered categories to quantify properties. These tools serve as proxies when direct physical measurement instruments are unavailable for assessing traits like strength or quality.
The authors suggest that the experimental measurement paradigm requires a more rigorous, theory-driven design process. This necessity arises because the model must satisfy specific requirements for additivity and invariance that statistical modeling does not always enforce.
The authors utilize a practical example to illustrate how different paradigms interpret the same data. This illustration serves to clarify the practical consequences of selecting one model over the other during instrument development.
The researchers measure the degree of a property, such as better or worse, using ordered categories. This measurement phenomenon relies on the assumption that responses reflect an underlying latent trait.
The authors claim that the choice of paradigm dictates the specific responsibilities of psychometricians. They propose that these experts must adapt their design strategies based on whether they prioritize statistical fit or structural invariance.
Related Concept Videos
Ratio Level of Measurement
A set of data measured using the ratio scale takes care of the ratio problem and provides complete information. Ratio scale data are like interval scale data, except they have a zero point and ratios can be calculated. For...
Ordinal Level of Measurement
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the...
Self-Report Tests of Personality
Nominal Level of Measurement
The data that cannot be measured but can be grouped into categories fall under the nominal level of measurement. Data that is measured using a nominal scale is...
Interval Level of Measurement
Data measured using the interval scale are similar to ordinal level data because they have a definite arrangement. However, in the interval level of measurement, the differences between data values are meaningful even though the data does not have a starting point.
Temperature is measured using the interval scale. It is measurable data, and the difference between the...
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
