Related Experiment Video
Updated: Dec 17, 2025

The Modular Design and Production of an Intelligent Robot Based on a Closed-Loop Control Strategy
Published on: October 14, 2017
A Multitasking-Oriented Robot Arm Motion Planning Scheme Based on Deep Reinforcement Learning and Twin
Chuzhao Liu1,2, Junyao Gao1,2, Yuanzhen Bi1,2
1Intelligent Robotics Institute, School of Mechatronical Engineering, Beijing Institute of Technology, 5 Nandajie, Zhongguancun, Haidian, Beijing 100081, China.
This study introduces a novel deep reinforcement learning (DRL) and digital twin approach for controlling humanoid robot arms. The method enables rapid, stable, and diverse multitasking for robots like the BHR-6, improving learning efficiency.
Area of Science:
- Robotics and Artificial Intelligence
- Humanoid Robot Control Systems
- Digital Twin Technology Applications
Background:
- Humanoid robots with arms are crucial for public acceptance and present significant robotics challenges.
- Digital twin technology aligns with Industry 4.0 and Made in China 2025 initiatives.
- Existing methods struggle with rapid, diverse, and stable motion planning for humanoid robot arms.
Purpose of the Study:
- To propose a combined deep reinforcement learning (DRL) and digital twin scheme for controlling humanoid robot arms.
- To develop a multitasking-oriented training approach for rapid and stable motion planning.
- To enhance DRL training efficiency by incorporating a priori knowledge and improving reward functions.
Main Methods:
- A Twin Synchro-Control (TSC) scheme integrating DRL with digital twin technology for robot arm control.
- Development of a data acquisition system to generate human joint angle data for training.
- Utilizing human joint angle data to refine the reward function of the Deep Deterministic Policy Gradient (DDPG) algorithm.
- Application to the BHR-6 humanoid robot model for simulation-based training.
Main Results:
- The proposed DRL with TSC scheme enables fast and diverse multitasking for humanoid robot arms.
- The approach effectively addresses the sparse reward problem in DRL using collected human motion data.
- Simulations demonstrated superior learning stability and convergence speed compared to vanilla DDPG.
- The trained humanoid robot successfully performed tasks not achievable with existing training methods.
Conclusions:
- The integration of DRL and digital twin technology offers a powerful solution for humanoid robot arm control.
- The TSC scheme with a priori knowledge significantly improves training efficiency and task performance.
- This method facilitates rapid adaptation and multi-task capability in complex humanoid robots.
Related Concept Videos
Planar Rigid-Body Motion
Planar motion is typically divided into three distinct categories. The first is rectilinear translation, demonstrated by a subway train that moves along...
One-Degree-of-Freedom System
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
Hierarchy of Motor Control
Muscle Coordination and Action
Agonists
Agonist muscles, often called prime movers, are the primary muscles responsible for producing a specific movement....
Three-Dimensional Force System:Problem Solving
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
Open and closed-loop control systems
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...

