Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs

Mohammad Shahab Sepehri1, Berk Tinaz1, Zalan Fabian1

  • 1Department of Electrical and Computer Engineering University of Southern California, Los Angeles, CA, USA.

Summary

Multimodal Large Language Models (MLLMs) struggle with mental visualization, the internal construction of visual patterns. A new benchmark, Hyperphantasia, reveals a significant performance gap compared to humans, highlighting an open challenge for AI.