Related Experiment Video
Updated: Jun 25, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Testing theory of mind in large language models and humans
James W A Strachan1, Dalila Albergo2,3, Giulia Borghini2
1Department of Neurology, University Medical Center Hamburg-Eppendorf, Hamburg, Germany. james.wa.strachan@gmail.com.
Large language models (LLMs) show human-like theory of mind abilities in some tasks, with GPT-4 matching or exceeding human performance in areas like false beliefs. However, both LLMs and humans struggle with detecting faux pas, indicating nuanced differences in artificial intelligence and human cognition.
Area of Science:
- Cognitive Science
- Artificial Intelligence
- Computational Linguistics
Background:
- Theory of Mind (ToM) is a key human cognitive ability, involving tracking others' mental states.
- Large Language Models (LLMs) are increasingly sophisticated, prompting comparisons with human cognitive functions.
- Debate exists on whether LLMs can exhibit human-like behavior in ToM tasks.
Purpose of the Study:
- To systematically compare human and LLM performance on a comprehensive suite of ToM tests.
- To investigate the capabilities of GPT and LLaMA2 families of LLMs in various ToM dimensions.
- To understand the nature of differences and similarities in ToM performance between humans and LLMs.
Main Methods:
- Administered a battery of ToM measurements to 1,907 human participants.
- Tested two LLM families (GPT and LLaMA2) repeatedly across the same ToM measures.
- Included tasks assessing false beliefs, indirect requests, irony, and faux pas detection.
Main Results:
- GPT-4 models performed at or above human levels on indirect requests, false beliefs, and misdirection.
- LLaMA2 outperformed humans only on faux pas detection, which was later found to be an illusory superiority.
- GPT-4's limitations stemmed from a conservative inference approach, not a fundamental lack of ability.
Conclusions:
- LLMs demonstrate behavior consistent with mentalistic inference, comparable to humans in certain ToM aspects.
- Systematic, nuanced testing is crucial for accurate comparisons between human and artificial intelligence.
- Differences in performance highlight the complexity of ToM and the need for further research into LLM cognition.
Related Concept Videos
Language and Cognition
Cognitivism
Previously dominated by behaviorism, which prioritized observable behaviors and largely ignored mental processes, psychology transformed in the 1950s. Cognitive psychologists argue that understanding how we think and process...
Observational Learning
Introduction to Cognitive Psychology
This field emerged in the mid-20th century, following a period dominated by behaviorism, which...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

