Related Experiment Videos
Why do reviewers disagree? Evidence from four funders and 134,000 reviews
Jan-Ole Hesselberg1,2, Pål Ulleberg1, Øystein Sørensen1
1Department of Psychology, University of Oslo, Oslo, Norway.
Abstract:
Grant peer review processes are pivotal in allocating substantial research funding, yet concerns about their reliability persist, primarily due to low inter-rater agreement. This study is a registered report that examines factors associated with agreement among peer reviewers in grant evaluations, leveraging data from 134,991 reviews across four Norwegian research funders. Using a cross-classified linear regression model, we explored the relationship between inter-rater agreement and multiple factors, including reviewer similarity, experience, expertise, research area, application characteristics, review depth, and temporal trends. Our analyses revealed several statistically significant associations, but most effects were small in absolute magnitude. Reviewer similarity in gender was associated with lower disagreement. The use of in-depth review procedures, referring to review roles with greater responsibility for evaluating the application, was also associated with slightly lower disagreement. In contrast, expert reviews, defined as assessments conducted by reviewers selected for subject-matter expertise rather than general panel membership, and certain research areas showed modestly higher disagreement. Application amount, reviewer age composition, and temporal trends exhibited limited influence. Only a small share of disagreement stemmed from stable reviewer tendencies (3.1 percent), whereas most variability arose at the application level (24.1 percent) and the residual level (55.2 percent), indicating that disagreement is largely application-specific rather than reviewer-driven. Inter-rater reliability was moderate across funders: average ICCs ranged from 0.50 to 0.65. Our findings only partly support previous research and indicate that disagreement is driven primarily by application-specific ambiguity and inherent variability in human judgment rather than systematic reviewer characteristics or program features. Consequently, interventions targeting reviewer behaviour alone are unlikely to substantially improve reliability; gains are more likely to come from clearer criteria, more structured review formats, and reducing application-level ambiguity, though such measures are expected to yield only incremental improvements.
Related Concept Videos
Social Proof
Confirmation Biases
Friedman Two-way Analysis of Variance by Ranks
Feedback Inhibition
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies