言語的類似性を介したスコアリング信頼性の再概念化
Ji Yoon Jung1, Ummugul Bezirhan1, Matthias von Davier1
1Boston College, Chestnut Hill, MA, USA.
Abstract:
Conventional cross-country scoring reliability in international large-scale assessments often depends on double scoring, which typically involves relatively small samples of multilingual responses. To extend the reach of reliability estimation, this study introduces the Linguistic-integrated Reliability Audit (LiRA), a novel method that measures scoring reliability using an entire dataset in a large-scale, multilingual context. LiRA automatically generates a second score for each response by analyzing its semantic alignment within a neighborhood of similar responses, then applies a weighted majority voting to determine a consensus score. Results demonstrate that LiRA provides a more comprehensive and systematic estimation of scoring reliability at the item, country, and language levels, while preserving the fundamental concepts of traditional reliability.
さらに関連する動画
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
関連する概念動画
Causes of Similarity-Dissimilarity Effect
Reliability and Validity
Kendall's Coefficient of Concordance
Spearman's Rank Correlation Test
Spearman's test calculates correlation by...
Language and Cognition
Improving Translational Accuracy
