Related Experiment Video
Updated: Jun 27, 2026

One Dimensional Turing-Like Handshake Test for Motor Intelligence
Published on: December 15, 2010
Human Shadows in Machine Minds: Quantitative Study Interpreting AI Responses to the Rorschach Test
Katalin Csigó1,2, György Cserey1
1Faculty of Information Technology and Bionics, Pázmány Péter Catholic University, Budapest, Budapest, Hungary.
Background:
Multimodal large language models (LLMs) can produce humanlike descriptions of images and emotionally colored dialogue, which motivates research on how psychological assessment methods might be adapted to evaluate model behavior under ambiguity. Projective tests such as the Rorschach inkblot test have rarely been applied to LLMs.
Objective:
This study assessed the feasibility of administering a full Rorschach protocol to multimodal LLMs and descriptively compared response features by using established Rorschach coding categories.
Methods:
We presented all 10 standard Rorschach cards to 3 multimodal LLMs (GPT-4o, Grok 3, and Gemini 2.0 Flash Thinking). We used the standard prompt ("What might it be?") and a prespecified fallback prompt for models that did not provide codable responses. We conducted an inquiry phase and coded responses using the Exner Comprehensive System, summarizing response count (R), location (W and D), determinants (eg, F, M, and C), and human-related content. As an exploratory step, we also prompted an additional LLM (Anthropic 3.7) to summarize and count response features and compared these outputs with manual tallies. For GPT-4o, we additionally tested image generation of its interpretations.
Results:
GPT-4o completed the administration using the standard prompt; Grok 3 and Gemini required the fallback prompt. The total number of responses was 15 for GPT-4o, 10 for Grok 3, and 20 for Gemini. GPT-4o and Grok 3 produced mainly whole-blot responses (13/15, 86.7% and 9/10, 90%, respectively), whereas Gemini produced mainly common-detail responses (16/20, 80%). Human movement determinants were more frequent in GPT-4o (7/15, 46.7%) and Grok 3 (3/10, 30%) than in Gemini (1/20, 5%). Human-themed contents occurred 46.7% (7/15), 50% (5/10), and 20% (4/20) of the time, respectively. Anthropic 3.7 reproduced some counts but showed errors in response and determinant tallies for 2 of the 3 models.
Conclusions:
Multimodal LLMs can generate Rorschach-like narratives that map onto standard coding categories, but outputs are sensitive to prompting and platform constraints and should not be interpreted as evidence of a model "inner world." LLM-assisted coding showed limitations. The emergent behavior of LLMs was examined using the Rorschach test, and their response phenotype, based on this analysis, showed deviations from typical human normative patterns. Future work should use controlled sampling, repeated administrations, and stimulus sets less likely to have been seen during training.
More Related Videos
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
08:01Virtual Hand with Ambiguous Movement between the Self and Other Origin: Sense of Ownership and 'Other-Produced' Agency
Published on: October 28, 2020
Related Concept Videos
Reason and Intuition
Introduction to Cognitive Psychology
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...