Related Experiment Video
Updated: Aug 5, 2026

Assessing Human Spatial Navigation in a Virtual Space and its Sensitivity to Exercise
Published on: January 26, 2024
SpatialPrompting: pose-aware keyframe prompting for 3D spatial QA toward smart indoor environments
Shun Taguchi1, Hideki Deguchi1, Takumi Hamazaki1
1Toyota Central R&D Labs., Inc., Nagakute, Aichi, Japan.
Abstract:
SpatialPrompting is a practical, training-free and model-agnostic framework for 3D spatial reasoning with off-the-shelf multimodal large language models, requiring no fine-tuning or 3D-specific inputs. It adopts a pose-aware, keyframe-driven prompting strategy: we select a compact, diverse set of frames using vision-language similarity, Mahalanobis distance, field of view, and image sharpness, and then verbalize each camera pose to enable multi-view, viewpoint-aware reasoning within a single structured prompt. Evaluations on ScanQA and SQA3D using GPT-4o and Gemini show that SpatialPrompting achieves competitive performance on ScanQA compared to existing approaches, while remaining slightly below the best fine-tuned methods on SQA3D under a training-free setting. On our Complex Spatial QA (CSQA) dataset, the proposed method improves accuracy from 62 to 78 (+16) over a GPT-4o baseline and consistently outperforms query-based keyframe selection. Furthermore, experiments across multiple models-including GPT-4o, Gemini, and Qwen-demonstrate that the framework generalizes across both proprietary and open-source models. By eliminating task-specific training and enabling reuse across different models without retraining, SpatialPrompting shifts the cost from model-specific training to flexible inference, offering a scalable and effective approach for real-world spatial reasoning as multimodal models continue to evolve.
Related Concept Videos
Depth Perception and Spatial Vision
Support Reactions in Three Dimensions
Ball and Socket Joint is one of the supports allowing free rotation about any axis. This freedom of rotation is...
Planar Rigid-Body Motion
Planar motion is typically divided into three distinct categories. The first is rectilinear translation, demonstrated by a subway train that moves along...
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
