导航多语言评估:测试开发,链接和评估的最佳实践.
Louise Badham1, María Elena Oliveri2, Stephan G Sireci3
1International Baccalaureate (United Kingdom).
Psicothema
|June 23, 2025
概括
在多种语言中创建评估提出了重大挑战. 本综述探讨了确保跨语言评估开发和验证的有效性和可比性的方法.
科学领域:
- 心理测量 心理测量 心理测量
- 教育测量教育的测量
- 应用语言学 应用语言学
背景情况:
- 多语言评估开发是复杂的,影响从项目创建到分数解释的所有阶段.
- 确保跨语言的得分可比性和有效性需要专门的方法.
研究的目的:
- 审查跨语言评估当前的文献和实践.
- 为从业人员提供基于研究的建议,用于多语言测试开发和验证.
主要方法:
- 对跨语言评估现有方法的文献综述.
- 分析测试开发,链接和验证中的实践.
主要成果:
- 从翻译转向同时开发多种语言的项目.
- 使用定量和定性方法进行跨语言评估.
- 方法的重点是将评估联系起来,并收集跨语言的有效性证据.
结论:
- 提供了当前多语言评估方法的概述.
- 为测试开发人员和研究人员提供实用建议.
- 强调在跨语言背景下有效性证据的重要性.
相关概念视频
Measures of Intelligence
7.9K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
7.9K
Language and Cognition
460
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
460
Reliability and Validity
13.2K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.2K
Improving Translational Accuracy
2.7K
2.7K
Multiple Comparison Tests
4.0K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.0K
Group Design
9.7K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
9.7K


