Related Experiment Video
Updated: Jul 13, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Validating a forced-choice method for eliciting quality-of-reasoning judgments
Alexandru Marcoci1, Margaret E Webb2, Luke Rowe3
1Centre for the Study of Existential Risk, University of Cambridge, 16 Mill Lane, Cambridge, CB2 1SB, UK. am3159@cam.ac.uk.
Forced-choice comparisons effectively assess written argument quality, with novices and experts accurately identifying better reasoning. This method offers a valid, reliable, and efficient approach for large-scale quality assessments.
Area of Science:
- Cognitive Psychology
- Educational Measurement
- Artificial Intelligence
Background:
- Evaluating the quality of written arguments is crucial but often subjective.
- Traditional methods can be time-consuming and require expert raters.
- Developing scalable and reliable assessment tools is an ongoing challenge.
Purpose of the Study:
- To investigate the criterion validity of forced-choice comparisons for assessing argument quality.
- To determine the reliability and efficiency of this assessment method for novices and experts.
- To explore methods for optimizing the assessment process and utilizing linguistic features.
Main Methods:
- Two studies were conducted using a forced-choice design where participants compared arguments.
- Inter-rater reliability was calculated to assess agreement between raters.
- Efficiency was improved using transitivity and an AVL tree method; a regression model analyzed linguistic features.
Main Results:
- Novices and experts accurately identified arguments supporting more accurate solutions (62.2% and 74.4% respectively).
- Participants correctly identified arguments from larger teams (82% for novices, 85% for experts) with high inter-rater reliability.
- Efficiency methods reduced judgment numbers with minimal accuracy loss; a regression model showed high correlation with objective scores.
Conclusions:
- Forced-choice comparisons provide a valid and reliable method for assessing reasoning quality, even for novices.
- The method is efficient and scalable for large-volume assessments.
- Leveraging transitivity and linguistic analysis offers further optimization for argument quality evaluation.
Related Concept Videos
Deductive Reasoning
For example, a researcher can deduce specific predictions...
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Reason and Intuition

