Related Experiment Video
Updated: Jun 4, 2026

07:29
A Novel Approach for Documenting Phosphenes Induced by Transcranial Magnetic Stimulation
Published on: April 1, 2010
Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
Mohammad Shahab Sepehri1, Berk Tinaz1, Zalan Fabian1
1Department of Electrical and Computer Engineering University of Southern California, Los Angeles, CA, USA.
Advances in Neural Information Processing Systems
|June 3, 2026
Summary
Multimodal Large Language Models (MLLMs) struggle with mental visualization, the internal construction of visual patterns. A new benchmark, Hyperphantasia, reveals a significant performance gap compared to humans, highlighting an open challenge for AI.
Area of Science:
- Cognitive Science
- Artificial Intelligence
- Computer Vision
Background:
- Mental visualization is crucial for human reasoning, prediction, and abstraction.
- Current benchmarks for Multimodal Large Language Models (MLLMs) focus on passive perception, not active internal visual construction.
- Assessing MLLMs' ability to internally construct visual patterns is vital for advancing AI capabilities.
Purpose of the Study:
- Introduce Hyperphantasia, a novel synthetic benchmark to evaluate MLLMs' mental visualization abilities.
- Provide a controlled environment to analyze MLLM performance across varying puzzle complexities.
- Bridge the gap between human cognitive skills and MLLM capabilities in visual reasoning.
Main Methods:
- Developed Hyperphantasia, a benchmark with four procedurally generated puzzles.
- Each puzzle is presented at three difficulty levels for granular performance analysis.
- Evaluated state-of-the-art MLLMs on the Hyperphantasia benchmark.
Main Results:
- A substantial performance gap exists between human capabilities and current MLLMs in mental visualization.
- Some MLLMs show partial ability in visual pattern recognition but lack robust internal visual construction.
- Reinforcement learning shows potential for enhancing visual simulation in MLLMs.
Conclusions:
- Robust mental visualization remains a significant challenge for current MLLMs.
- The Hyperphantasia benchmark provides a valuable tool for future research and development in AI cognition.
- Publicly available dataset and code facilitate further investigation into MLLM visual reasoning.

