Related Experiment Video
Updated: Jul 14, 2026

10:58
Multimedia Battery for Assessment of Cognitive and Basic Skills in Mathematics (BM-PROMA)
Published on: August 28, 2021
Using reliability generalization methods to explore measurement error: an illustration using the MMPI-2 PSY-5 scales
1Social Sciences Division, Pepperdine University, Malibu, CA 90263, USA. steve.rouse@pepperdine.edu
Journal of Personality Assessment
|May 24, 2007
Summary
Reliability generalization (RG) reveals that test score reliability varies by sample, not just the test itself. Researchers should calculate reliability for their specific samples rather than relying on published coefficients.
Area of Science:
- Psychological assessment
- Psychometrics
- Meta-analysis
Background:
- Test score reliability is often assumed to be constant.
- However, reliability can vary significantly across different samples.
- Reliability generalization (RG) provides a method to systematically study this variation.
Purpose of the Study:
- To demonstrate the application of RG analysis to the MMPI-2 Personality Psychopathology 5 (PP5) scales.
- To examine the sample dependency of score reliability for the PP5 scales.
- To identify factors influencing score reliability variations.
Main Methods:
- A meta-analytic reliability generalization (RG) study was conducted.
- Sixty-three reliability coefficients for each of the MMPI-2 PP5 scales were collected and analyzed.
- Statistical analyses examined mean reliability differences and relationships with sample characteristics.
Main Results:
- Significant variability in reliability coefficients across samples was observed, supporting the sample-dependent nature of reliability.
- Reliability differed significantly across the five PP5 scales, with Negative Emotionality being most reliable and Aggression/Disconstraint least reliable.
- Score reliability was lower in nonclinical settings compared to clinical settings for some scales; sample sex composition and size were not significant predictors.
Conclusions:
- The study underscores the critical need for researchers to compute reliability estimates specific to their own research samples.
- RG analysis offers valuable insights into the psychometric properties of personality tests and informs their appropriate use.
- Findings highlight the importance of considering sample characteristics, such as clinical vs. nonclinical settings, when interpreting test score reliability.
Related Concept Videos
Self-Report Tests of Personality
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
Reliability and Validity
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Diagnostic and Statistical Manual of Mental Disorders (DSM)
The Diagnostic and Statistical Manual of Mental Disorders (DSM) serves as the primary classification system for mental health disorders, providing standardized diagnostic criteria for clinicians and researchers. First published by the American Psychiatric Association (APA) in 1952, the DSM has undergone several revisions to reflect evolving psychiatric understanding. The fifth edition, DSM-5, released in 2013, introduced key updates that expanded diagnostic categories and modified diagnostic...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...