对监管绩效评级的评级者间可靠性的更新的元分析
You Zhou1, Paul R Sackett1, Winny Shen2
1Department of Psychology, University of Minnesota, Twin Cities.
The Journal of applied psychology
|January 25, 2024
概括
这次元分析发现,使用一种新的方法,对工作绩效评级的评级者间可靠性更高 (r = .65). 可靠性因工作类型而异,这表明特定于工作的估计对于准确的校正至关重要.
科学领域:
- 组织心理学 组织心理学
- 人力资源管理 人力资源管理
- 工业组织心理学 工业组织心理学
背景情况:
- 工作绩效是组织研究的核心.
- 准确地衡量工作绩效是至关重要的.
- 监管评级的Interrater可靠性是常用的,但需要更新分析.
研究的目的:
- 对监督性工作绩效评级的评级者之间的可靠性进行更新的元分析.
- 检查影响评级者间可靠性的因素,如工作复杂性和管理层级.
- 将发现与以前的元分析进行比较,并完善方法论方法.
主要方法:
- 使用了一种新的元分析程序 (莫里斯估计器),将研究内部和研究间的差异纳入其中.
- 分析了132个独立的监督性工作绩效评级样本.
- 调查了职位复杂度,管理层级,评级目的,绩效指标和评级者视角的影响.
主要成果:
- 与之前的元分析相比,实现了更高的interrater可靠性估计 (r = .65).
- 证实了评价者之间的可靠性因工作类型而异 (管理职位的r = .57与非管理职位的r = .68).
- 莫里斯估计器阻止了大样本研究不成比例地影响结果.
结论:
- 建议不要使用单一的整体平均值来衡量评级者之间的可靠性.
- 建议使用特定职位或本地可靠性估计进行减弱校正.
- 强调在评估绩效评级可靠性时考虑职位特征的重要性.
更多相关视频
09:16Use of a Video Scoring Anchor for Rapid Serial Assessment of Social Communication in Toddlers
Published on: March 14, 2018
10.3K
08:40Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
Published on: June 12, 2019
7.5K
相关概念视频
Friedman Two-way Analysis of Variance by Ranks
197
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
197
Kendall's Coefficient of Concordance
339
Kendall's Coefficient of Concordance (W), also known as Kendall's W, is a non-parametric statistical measure used to assess the agreement or concordance between multiple raters or judges when they rank a set of items. It is often used when you have ordinal data (ranks) and you want to see if there is consistency or consensus among the raters. It is widely applied in research areas such as psychology, medicine, and social sciences, where multiple judges are asked to rank or rate subjects...
339
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Self-Discrepancy Theory
18.3K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.
18.3K
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K
Self-Report Tests of Personality
352
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
352
