Related Experiment Video
Updated: Jul 16, 2025

A Cross-Disciplinary and Multi-Modal Experimental Design for Studying Near-Real-Time Authentic Examination Experiences
Published on: September 4, 2019
Hawks and Doves: Perceptions and Reality of Faculty Evaluations
Jillian Zavodnick1, Jonathan Doroshow2, Sarah Rosenberg1
1Sidney Kimmel Medical College, Thomas Jefferson University, Philadelphia, USA.
Objectives:
Internal medicine clerkship grades are important for residency selection, but inconsistencies between evaluator ratings threaten their ability to accurately represent student performance and perceived fairness. Clerkship grading committees are recommended as best practice, but the mechanisms by which they promote accuracy and fairness are not certain. The ability of a committee to reliably assess and account for grading stringency of individual evaluators has not been previously studied.
Methods:
This is a retrospective analysis of evaluations completed by faculty considered to be stringent, lenient, or neutral graders by members of a grading committee of a single medical college. Faculty evaluations were assessed for differences in ratings on individual skills and recommendations for final grade between perceived stringency categories. Logistic regression was used to determine if actual assigned ratings varied based on perceived faculty's grading stringency category.
Results:
"Easy graders" consistently had the highest probability of awarding an above-average rating, and "hard graders" consistently had the lowest probability of awarding an above-average rating, though this finding only reached statistical significance only for 2 of 8 questions on the evaluation form (P = .033 and P = .001). Odds ratios of assigning a higher final suggested grade followed the expected pattern (higher for "easy" and "neutral" compared to "hard," higher for "easy" compared to "neutral") but did not reach statistical significance.
Conclusions:
Perceived differences in faculty grading stringency have basis in reality for clerkship evaluation elements. However, final grades recommended by faculty perceived as "stringent" or "lenient" did not differ. Perceptions of "hawks" and "doves" are not just lore but may not have implications for students' final grades. Continued research to describe the "hawk and dove effect" will be crucial to enable assessment of local grading variation and empower local educational leadership to correct, but not overcorrect, for this effect to maintain fairness in student evaluations.
More Related Videos
07:32Use of Galvanic Skin Responses, Salivary Biomarkers, and Self-reports to Assess Undergraduate Student Performance During a Laboratory Exam Activity
Published on: February 10, 2016
04:12Mixed Reality for Education MRE Implementation and Results in Online Classes for Engineering
Published on: June 23, 2023
Related Concept Videos
Self-Evaluation: Self-Enhancement and Self-Verification
Surveys
Factors Affecting Perception
An illustrative example of a perceptual set is the scenario where an airline pilot told...
Nursing Evaluation
The Sense of Self: Reflected Self-Appraisal and Social Comparison
The Representativeness Heuristic