Related Experiment Video
Updated: Jan 10, 2026

A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
ChatGPT does not replicate human moral judgments: the importance of examining metrics beyond correlation to assess
Matthew Grizzard1, Rebecca Frazer2, Andrew Luttrell3
1School of Communication, The Ohio State University, Columbus, Ohio, USA. grizzard.6@osu.edu.
Abstract:
The rise of generative artificial intelligence has prompted claims that large language models (LLMs) can substitute for human participants, particularly in moral judgment tasks where correlations between ChatGPT and humans approach r = 1.00. In response, we conducted a pre-registered study where two LLMs (text-davinci-003 and GPT-4o) predicted human moral judgments of 60 scenarios prior to a large human sample (N = 940) rating them. Despite strong correlations, difference scores revealed substantial, systematic errors: Compared to humans, LLMs provided more extreme morality ratings of moral and neutral scenarios and more extreme immorality ratings of immoral ones. Moreover, ChatGPT differed significantly and with moderate to large effect sizes from human averages on ~ 87% of scenarios. Further, LLM ratings clustered around a restricted number of values, failing to reflect human variability. Re-examination of earlier published data also reflected this clumping. We conclude that broader evaluation criteria are needed for comparing LLM predictions and human responses in moral reasoning tasks.
More Related Videos
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
09:38Generalized Psychophysiological Interaction PPI Analysis of Memory Related Connectivity in Individuals at Genetic Risk for Alzheimer's Disease
Published on: November 14, 2017
Related Concept Videos
Detection of Gross Error: The Q Test
Quantifying and Rejecting Outliers: The Grubbs Test
Regression Toward the Mean
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Reliability and Validity
Fundamental Attribution Error