Related Experiment Video
Updated: Oct 5, 2026

Using Virtual Reality to Transfer Motor Skill Knowledge from One Hand to Another
Published on: September 18, 2017
WIS: world-guided instruction semantic diffusion policy for predictive bimanual manipulation
Xukun Liu1, Liguo Hou1, Pu Cheng1
1Northwest Institute of Mechanical and Electrical Engineering, Xianyang, Shaanxi, China.
Introduction:
Robotic bimanual manipulation poses significant challenges for reactive visuomotor policies, which inherently lack the capacity to anticipate future environmental dynamics, especially in long-horizon and semantically complex tasks. To address this limitation, there is a critical need for frameworks that integrate predictive modeling with task-level reasoning to enable coordinated, forward-looking dual-arm control.
Methods:
We introduce the World-guided Instruction Semantic Diffusion Policy (WIS), a novel hierarchical framework that tightly integrates a language-driven world model with a diffusion-based policy for dual-arm coordination. The architecture comprises three core modules: (1) a multimodal language-visual perception module that fuses 3D point cloud features with natural language embeddings via cross-modal attention for task-relevant semantic grounding; (2) a language-conditioned world model that explicitly forecasts future state trajectories conditioned on current states, action sequences, and language instructions, providing forward-looking predictive reasoning; and (3) a hierarchical diffusion policy with a high-level latent planner and a low-level diffusion controller, wherein the world model's predictions actively guide action refinement through a prediction-aware sampling mechanism. This design cleanly separates high-level task reasoning from low-level motor generation while ensuring every action remains grounded in both immediate observations and anticipated future outcomes. We extensively evaluated WIS on the RoboTwin 2.0 platform using Aloha-AgileX bimanual robots across multiple challenging manipulation tasks.
Results:
Experimental results on the RoboTwin 2.0 benchmark demonstrate that WIS consistently outperforms strong baselines, achieving superior success rates across all evaluated tasks.
Discussion:
The proposed method offers an efficient and robust solution for temporally coherent bimanual manipulation, with minimal computational overhead relative to baseline approaches. These findings underscore the value of integrating predictive world models with semantic instructions to enhance both the foresight and dexterity of robotic dual-arm systems in complex, real-world settings.

