Related Experiment Video
Updated: Jun 30, 2025

08:12
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
2.5K
Performance evaluation of automated scoring for the descriptive similarity response task
Ryunosuke Oka1, Takashi Kusumi2, Akira Utsumi3
1Mitsubishi Electric Corporation, Kamakura, Kanagawa, 247-8501, Japan. Qualia1006@gmail.com.
Scientific Reports
|March 15, 2024
Summary
A machine learning model achieved human-level performance in scoring the Japanese Semantic Similarity Test (SST). The automated system demonstrated high accuracy and inter-rater reliability comparable to human scorers.
Area of Science:
- Natural Language Processing
- Artificial Intelligence in Education
- Psychometrics
Background:
- The Semantic Similarity Test (SST) assesses understanding of conceptual relationships.
- Automated scoring systems are increasingly used in educational and psychological assessments.
- Evaluating the performance of machine learning models against human judgment is crucial for their adoption.
Purpose of the Study:
- To investigate the efficacy of a machine learning model in replicating human scoring on the Japanese SST.
- To compare the accuracy and reliability of automated scoring with human performance.
Main Methods:
- A bidirectional encoder representations from transformer (BERT) model was trained using participant responses and classification criteria from the SST.
- Fivefold cross-validation was employed for model evaluation.
- Model performance was assessed using accuracy and inter-rater reliability metrics.
Main Results:
- The model achieved 83% accuracy when trained on both responses and classification criteria (Experiment 1).
- In a separate test set, the model scored 80% accuracy (Experiment 2).
- Human-model inter-rater reliability (0.63) closely matched human-human reliability (0.67-0.72).
Conclusions:
- Machine learning models can achieve human-level performance in scoring the Japanese SST.
- Automated scoring systems offer a reliable alternative to human scoring for similarity tasks.
- This research supports the use of AI for objective and efficient assessment in language and cognitive studies.
Related Concept Videos
Response Surface Methodology
129
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
129
Wilcoxon Signed-Ranks Test for Matched Pairs
123
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
123
Spearman's Rank Correlation Test
783
Spearman's rank correlation test, also known as Spearman's rho, is a nonparametric method for assessing the strength and direction of association between two variables. This test is particularly valuable when the data distribution is unknown or when the assumption of normality does not hold. Named after the English psychologist and statistician Dr. Charles Edward Spearman, it serves as the nonparametric counterpart to Pearson's correlation coefficient.
Spearman's test calculates...
Spearman's test calculates...
783
Review and Preview
7.4K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
7.4K

