Related Experiment Videos
Converging on the tipping point: a diagnostic methodology for standard setting
1Pearson Vue, 1 N. Dearborn St., Suite 1000, Chicago, IL 60602, USA. John.Stahl@pearson.com
Summary
This study evaluates the Angoff and Bookmark standard-setting procedures, proposing an enhanced method with diagnostic indices for improved educational assessment.
Area of Science:
- Educational Measurement
- Psychometrics
Background:
- The Angoff and Bookmark procedures are widely used for setting performance standards in educational and professional assessments.
- Both methods have documented strengths and weaknesses that can impact the reliability and validity of standard-setting outcomes.
Purpose of the Study:
- To critically analyze the limitations of the Angoff and Bookmark standard-setting procedures.
- To introduce and evaluate an alternative standard-setting approach incorporating diagnostic indices.
Main Methods:
- A comparative analysis of the Angoff and Bookmark procedures was conducted.
- A novel standard-setting methodology was developed, integrating three diagnostic indices.
- The proposed method was applied to three distinct standard-setting datasets.
Main Results:
- The study identified specific weaknesses in the traditional Angoff and Bookmark methods.
- The alternative approach demonstrated potential for enhancing the diagnostic utility of standard-setting.
- Results from the three datasets indicated the practical applicability of the new procedure.
Conclusions:
- The proposed alternative standard-setting approach offers a valuable enhancement over existing methods.
- Diagnostic indices can provide deeper insights into the standard-setting process.
- Further research is warranted to validate and refine this enhanced methodology.
Related Concept Videos
Decision Making: Traditional Method
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Testing a Claim about Standard Deviation
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Significance Testing: Overview
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
Regression Toward the Mean
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when researchers try to extrapolate results...
Empirical Method to Interpret Standard Deviation
The empirical rule, also known as the three-sigma rule, allows a statistician to interpret the standard deviation in a normally distributed dataset. The rule states that 68% of the data lies within one standard deviation from the mean, 95% lies within two standard deviations from the mean, and 99.7% lies within three standard deviations from the mean. Additionally, this rule is also called the 68-95-99.7 rule.
This rule is used widely in statistics to calculate the proportion of data values...
This rule is used widely in statistics to calculate the proportion of data values...