Related Experiment Video
Updated: Jun 7, 2025

Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
Testing AI on language comprehension tasks reveals insensitivity to underlying meaning
Vittoria Dentella1,2, Fritz Günther3, Elliot Murphy4
1Universitat Rovira i Virgili, Tarragona, Spain. vittoria.dentella@unipv.it.
Large Language Models (LLMs) demonstrate chance accuracy on language comprehension tasks, significantly underperforming human performance. Current AI models lack human-like linguistic understanding due to potential deficiencies in compositional reasoning.
Area of Science:
- Artificial Intelligence
- Computational Linguistics
- Cognitive Science
Background:
- Large Language Models (LLMs) are increasingly used in diverse applications, leading to claims of human-like linguistic capabilities.
- Moravec's Paradox highlights the difficulty of replicating human-like skills in AI, even for seemingly simple tasks.
Purpose of the Study:
- To systematically evaluate the language comprehension and reasoning abilities of state-of-the-art LLMs.
- To compare LLM performance against human baselines on a novel benchmark assessing linguistic understanding.
Main Methods:
- Seven state-of-the-art LLMs were tested on a benchmark of comprehension questions.
- Models were prompted in two settings (one-word and open-length replies) across a dataset of 26,680 data points.
- A human baseline was established by testing 400 individuals on the same prompts.
Main Results:
- LLMs performed at chance accuracy and exhibited considerable variability in their responses.
- Human participants significantly outperformed all tested LLMs in quantitative accuracy.
- Qualitative analysis revealed non-human errors in LLM language understanding.
Conclusions:
- Current LLMs, despite their utility, do not possess human-like language understanding.
- The lack of a compositional operator for grammatical and semantic information may explain LLMs' limitations.
- Further research is needed to bridge the gap between AI and human linguistic capabilities.
More Related Videos
08:32Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
10:11Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism
Published on: December 14, 2012
Related Concept Videos
Language and Cognition
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Intelligence
Stereotypes, Prejudice, and Discrimination
Fundamental Attribution Error
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...