Related Experiment Video
Updated: Nov 7, 2025

19:44
A Tactile Automated Passive-Finger Stimulator TAPS
Published on: June 3, 2009
13.9K
Certainty-Based Marking on Multiple-Choice Items: Psychometrics Meets Decision Theory
Qian Wu1, Monique Vanerum2, Anouk Agten2
1Center for Educational Effectiveness and Evaluation, KU Leuven, Dekenstraat 2, bus 3773, 3000, Leuven, Belgium.
Psychometrika
|April 30, 2021
Summary
Certainty-based marking (CBM) helps assess knowledge but can be influenced by risk attitudes. Students often under-report certainty when confident and over-report when uncertain, affecting test scoring.
Area of Science:
- Educational Psychology
- Psychometrics
- Decision Science
Background:
- Traditional multiple-choice tests cannot distinguish knowledge from uncertainty.
- Certainty-based marking (CBM) incorporates examinee certainty into scoring.
- Prospect theory suggests risk attitudes influence decision-making, potentially impacting CBM.
Purpose of the Study:
- To investigate response behaviors in Certainty-based marking (CBM) among physiotherapy students.
- To integrate psychometric modeling with decision-making theory to analyze CBM.
- To understand how risk attitudes affect certainty reporting in a CBM context.
Main Methods:
- Utilized item response theory to model the probability of correct responses.
- Applied cumulative prospect theory to estimate examinee risk attitudes.
- Conducted a case study with 334 first-year physiotherapy students across six CBM examinations.
Main Results:
- Student certainty levels in CBM were significantly influenced by their risk attitudes.
- Risk-averse and loss-averse students tended to under-report certainty with high success probabilities.
- Risk-seeking students tended to over-report certainty with low success probabilities.
Conclusions:
- Risk attitudes play a crucial role in how students report certainty under CBM scoring.
- The findings highlight a potential bias in CBM due to non-rational decision-making.
- Understanding these behaviors is key to refining CBM for more accurate assessment.
Keywords:
IRTcertainty-based markingcumulative prospect theoryhierarchical Bayesian estimationmultiple-choice questionsMore Related Videos
Related Concept Videos
Decision Making: P-value Method
6.2K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.2K
Decision Making
398
Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Automatic decision-making is fast, intuitive, and relies on gut feelings...
398
Decision Making: Traditional Method
4.6K
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
4.6K
Multiple Comparison Tests
4.1K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.1K
Measures of Intelligence
8.0K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.0K
Self-Report Tests of Personality
540
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
540

