Related Experiment Video
Updated: Aug 6, 2026

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
Deep Reinforcement Learning for Real-World Humanoid Robot Locomotion Control with Automatic Reward Learning
Renzhi Lu1, Jie Wang2, Zonghe Shao2
1School of Artificial Intelligence and Automation, Key Laboratory of Image Processing and Intelligent Control, Engineering Research Center of Autonomous Intelligent Unmanned Systems, Chinese Ministry of Education, Huazhong University of Science and Technology, Wuhan 430074, China.
This study introduces an automatic reward learning method for humanoid robot locomotion control using deep reinforcement learning (DRL). The approach enhances learning efficiency and performance, enabling successful sim-to-real transfer for practical applications.
Area of Science:
- Robotics
- Artificial Intelligence
- Control Systems
Background:
- Humanoid robots are complex systems requiring advanced motion control for diverse applications.
- Deep Reinforcement Learning (DRL) offers adaptive control but faces challenges in reward function design for real-world tasks.
- Manual reward engineering for DRL is time-consuming, requires expertise, and can lead to suboptimal performance or mission failure.
Purpose of the Study:
- To develop an automatic reward learning method for DRL-based humanoid robot locomotion control.
- To address the limitations of manual reward function design in DRL.
- To improve the efficiency and robustness of humanoid robot motion control.
Main Methods:
- A bilevel optimization framework was developed for automatic reward learning.
- The upper level adaptively constructs and optimizes the reward function.
- The lower level employs DRL to learn the locomotion control policy using the learned reward function.
Main Results:
- The proposed method significantly improved learning efficiency and performance compared to manual reward functions.
- Experiments demonstrated effectiveness in simulation (MuJoCo, Isaac Lab) and real-world deployment (Unitree G1 robot).
- Successful sim-to-real transfer of control policies was achieved, enhancing deployment capabilities.
Conclusions:
- Automatic reward learning offers a more efficient and effective approach for DRL in humanoid robot locomotion.
- The method accelerates the development and deployment of stable, adaptive humanoid robots.
- This work provides a promising pathway for advancing practical humanoid robot applications.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Reinforcement Schedules
Once a behavior is learned,...
