Related Experiment Videos
Why do experts miss AI's errors? Evidence from a randomized labeling experiment
Sofoklis Goulas1,2,3,4, Rigissa Megalokonomou4,5,6,7, Panagiotis Sotirakopoulos8,9
1Research Team, foundry10, 4244 University Way NE, #85538, Seattle, WA 98145, USA.
PNAS Nexus
|June 11, 2026
Summary
Human experts are less likely to correct harsh AI-generated scores than human ones. This deference to artificial intelligence (AI) depends on error direction and perceived AI credibility, impacting AI-human collaboration.
Area of Science:
- Human-Computer Interaction
- Educational Psychology
- Artificial Intelligence Ethics
Background:
- Organizations increasingly use algorithmic decision aids, necessitating human oversight to mitigate automated errors.
- Understanding when human experts override or accept algorithmic recommendations is crucial for accountability.
Purpose of the Study:
- To investigate how the source (human vs. AI) and direction (harsh vs. lenient) of inaccurate algorithmic scores influence expert correction behavior.
- To identify the factors mediating expert deference to artificial intelligence (AI) in educational grading.
Main Methods:
- A preregistered randomized experiment involving education experts reviewing student work with inaccurate, source-labeled scores (human or AI-generated).
- Independent variation of score inaccuracy as either too harsh or too lenient.
- Measurement of the grading fairness gap as the primary outcome, assessing expert correction accuracy.
Main Results:
- A significantly larger grading fairness gap (22%) occurred when harsh AI-generated scores were reviewed compared to human-generated scores.
- In cases of lenient scores, the grading fairness gap was statistically indistinguishable between AI and human labels.
- Mediation analysis indicated perceived AI ability and responsibility explained over half the effect in harsh scenarios, with weaker attributions in lenient scenarios.
Conclusions:
- Expert deference to AI is contingent on the direction of algorithmic errors and the credibility signaled by the AI, not solely on automation.
- Findings offer design insights for fostering accountable human-AI collaboration, emphasizing the importance of error context in AI oversight.
Related Concept Videos
Fundamental Attribution Error
According to some social psychologists, people tend to overemphasize internal factors as explanations—or attributions—for the behavior of other people. They tend to assume that the behavior of another person is a trait of that person, and to underestimate the power of the situation on the behavior of others. They tend to fail to recognize when the behavior of another is due to situational variables, and thus to the person’s state. This erroneous assumption is called the fundamental attribution...
Cause and Effect
While variables are sometimes correlated because one does cause the other, it could also be that some other factor, a confounding variable, is actually causing the systematic movement in our variables of interest. For instance, as sales in ice cream increase, so does the overall rate of crime. Is it possible that indulging in your favorite flavor of ice cream could send you on a crime spree? Or, after committing crime do you think you might decide to treat yourself to a cone?
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
Confirmation Biases
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
Detection of Gross Error: The Q Test
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...