Reporting Valid and Reliable Overall Scores and Domain Scores Using Bi-Factor Model
Yue Liu1, Zhen Li2, Hongyun Liu1
1Beijing Normal University, China.
Applied Psychological Measurement
|September 20, 2019
Summary
This study introduces six methods for reporting scores from bi-factor models, finding that Bifactor-M4 and Bifactor-M6 offer the most accurate and reliable overall and domain scores, especially in complex testing scenarios.
Area of Science:
- Psychometrics
- Educational Measurement
- Statistical Modeling
Background:
- Increasing demand for accurate diagnostic information in large-scale testing programs.
- Limited research on reliable reporting of total and domain scores using bi-factor models.
Purpose of the Study:
- To introduce and compare six novel methods for reporting overall and domain scores derived from bi-factor models.
- To evaluate the performance of these methods against Yao's multidimensional item response theory (MIRT) method.
Main Methods:
- Development of six weighted composite score reporting methods for bi-factor models.
- Comparison using simulated data (varying test length, dimensionality, correlation, sample size) and empirical data.
- Evaluation metrics included score accuracy, reliability, and parameter recovery.
Main Results:
- Bifactor-M4 and Bifactor-M6 demonstrated superior accuracy and reliability for overall and domain scores across most conditions.
- These methods excelled particularly with longer tests, higher dimensional correlations, and more dimensions.
- Bifactor-M4 showed the best recovery of true ability parameters; Bifactor-M2 performed poorly on overall scores; Bifactor-M1 yielded the worst estimations.
Conclusions:
- Bifactor-M4 and Bifactor-M6 are recommended for reporting scores from bi-factor models due to their accuracy and reliability.
- The choice of weighting method significantly impacts score estimation quality.
- Further research should explore the practical implementation and interpretation of these advanced scoring methods.
Related Concept Videos
Self-Report Tests of Personality
782
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
782
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Measures of Intelligence
8.3K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.3K
Factorial Design
13.7K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.7K
Confidence Coefficient
10.5K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
10.5K
Self-Evaluation Maintenance Model
287
The Self-Evaluation Maintenance (SEM) model offers a psychological framework to understand how individuals’ self-esteem is influenced by the achievements of others, particularly those with whom they share close personal bonds. The SEM model operates when personal rather than social identity guides individuals. Central to this model is the notion that individuals have an inherent desire to preserve a favorable self-image, which is continuously shaped by interpersonal comparisons and...
287


