Related Experiment Video
Updated: Jun 30, 2025

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
Performance evaluation of automated scoring for the descriptive similarity response task
Ryunosuke Oka1, Takashi Kusumi2, Akira Utsumi3
1Mitsubishi Electric Corporation, Kamakura, Kanagawa, 247-8501, Japan. Qualia1006@gmail.com.
Abstract:
We examined whether a machine-learning-based automated scoring system can mimic the human similarity task performance. We trained a bidirectional encoder representations from transformer-model based on the semantic similarity test (SST), which presented participants with a word pair and asked them to write about how the two concepts were similar. In Experiment 1, based on the fivefold cross validation, we showed the model trained on the combination of the responses (N = 1600) and classification criteria (which is the rubric of the SST; N = 616) scored the correct labels with 83% accuracy. In Experiment 2, using the test data obtained from different participants in different timing from Experiment 1, we showed the models trained on the responses alone and the combination of responses and classification criteria scored the correct labels in 80% accuracy. In addition, human-model scoring showed inter-rater reliability of 0.63, which was almost the same as that of human-human scoring (0.67 to 0.72). These results suggest that the machine learning model can reach human-level performance in scoring the Japanese version of the SST.
Related Concept Videos
Response Surface Methodology
The process of RSM involves several key steps:
Wilcoxon Signed-Ranks Test for Matched Pairs
Spearman's Rank Correlation Test
Spearman's test calculates...
Review and Preview
Percentiles are a type of fractile that partition data into...

