Related Experiment Video
Updated: Aug 14, 2026

05:28
Application of a Dual Upper Limb Task-Oriented Robotic System for the Functional Recovery of the Upper Limb in Stroke Patients
Published on: October 11, 2024
Language-Guided Dual-Mode Policy for Dual-Arm Manipulation
Jianghao Sun1, Pengjun Mao1, Yu Wang1
1School of Mechanical and Electrical Engineering, Henan University of Science and Technology, Luoyang 471000, China.
Sensors (Basel, Switzerland)
|August 13, 2026
Summary
We introduce a novel Language-Guided Dual-Mode Policy (LGDM) to improve dual-arm robot coordination. LGDM effectively decouples synchronous and asynchronous tasks, enhancing manipulation performance and adaptability to new instructions.
Area of Science:
- Robotics and Artificial Intelligence
- Machine Learning for Manipulation
Background:
- Dual-arm robot manipulation tasks exhibit complex synchronous coordination or asynchronous sequential execution patterns.
- Existing action-generative policies often use single-network architectures, failing to address distinct temporal dependencies and action distributions, leading to cross-mode interference.
Purpose of the Study:
- To propose a novel framework, the Language-Guided Dual-Mode Policy (LGDM), to explicitly decouple synchronous and asynchronous dual-arm motion modes.
- To enhance robot manipulation by addressing cross-mode interference in multi-task learning.
Main Methods:
- LGDM utilizes shared language embeddings to construct parallel branches for language routing and visual perception.
- A language mode routing layer predicts motion modes, guiding an execution layer with independent synchronous and asynchronous policy branches.
- FiLM (Feature-wise Linear Modulation) is employed for task-aware dynamic feature modulation in the vision perception layer.
Main Results:
- LGDM demonstrated superior performance on the RoboTwin 2 benchmark, outperforming existing baselines with an absolute improvement of 14.5 percentage points over RDT.
- The framework achieved robust execution performance even with unseen temporal-coordination instruction variants at inference time.
Conclusions:
- The proposed LGDM framework effectively decouples dual-arm robot motion modes, leading to significant performance gains.
- LGDM offers enhanced adaptability and robustness for complex dual-arm manipulation tasks under language guidance.
