Related Experiment Video
Updated: May 1, 2026

Photorealistic Learned Landscapes for Augmented Reality
Published on: June 27, 2025
A retrieval-augmented framework enabling VLM spatial awareness for object-centric robot manipulation
Kai Chen1, Chengkun Li1, Chang Tu1
1Department of Computer Science and Engineering, Chinese University of Hong Kong, HKSAR, China.
Retrieval-Augmented Manipulation (RAM) enables vision-language models to perform precise robotic tasks by grounding language in 3D object representations. This framework bridges semantic understanding and geometric execution for enhanced robot intelligence.
Area of Science:
- Robotics
- Artificial Intelligence
- Computer Vision
Background:
- Vision-language models (VLMs) struggle with precise spatial reasoning for robotic manipulation.
- Existing VLMs lack the intrinsic spatial intelligence for object placement and orientation tasks.
Purpose of the Study:
- Introduce Retrieval-Augmented Manipulation (RAM) to bridge the semantic-to-geometric gap in robotic manipulation.
- Equip general-purpose vision foundation models with spatial reasoning capabilities for complex tasks.
Main Methods:
- Developed an object-centric framework (RAM) grounding abstract concepts into 3D representations.
- Augmented VLMs with grounded 3D information to decompose instructions into precise subgoals.
- Utilized a real-world robot for zero-shot execution of manipulation tasks.
Main Results:
- RAM successfully executed complex spatial language instructions in a zero-shot setting.
- Demonstrated spatially aware manipulation from a single 2D image and adaptive replanning.
- Validated generalization to unseen objects and robustness to shape variations and occlusions on the CO3D dataset.
Conclusions:
- RAM provides a structured bridge between semantic intent and geometric execution for robotic systems.
- This framework is a critical step toward developing more physically intelligent and general-purpose robots.
- The object-centric approach enhances VLM spatial reasoning for real-world manipulation challenges.
More Related Videos
Related Concept Videos
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Virtual Work for a System of Connected Rigid Bodies
Next,...
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Three-Dimensional Force System:Problem Solving
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
Depth Perception and Spatial Vision

