Related Experiment Video
Updated: Jan 10, 2026

SSVEP-based Experimental Procedure for Brain-Robot Interaction with Humanoid Robots
Published on: November 24, 2015
A Two-Stage Reinforcement Learning Framework for Humanoid Robot Sitting and Standing-Up
Xisheng Jiang1,2,3,4, Shihai Zhao1,2, Yudi Zhu1,2,3,4
1School of Optoelectronic Information and Computer Engineering, University of Shanghai for Science and Technology, Shanghai 200093, China.
Abstract:
In human daily-life scenarios, humanoid robots need not only to stand up smoothly but also to autonomously sit down for rest, energy management, and interaction. This capability is crucial for enhancing their autonomy and practicality. However, both sitting and standing involve complex dynamics constraints, diverse initial postures, and unstructured terrains, which make traditional hand-crafted controllers insufficient for multi-scenario demands. Reinforcement Learning (RL), with its generalization ability across high-dimensional state spaces and complex tasks, offers a promising solution for automatically generating motion control policies. Nevertheless, policies trained directly with RL often produce abrupt motions, making it difficult to balance smoothness and stability. To address these challenges, we propose a two-stage reinforcement learning framework: In the first stage, we focus on exploration and train initial policies for both sitting and standing, with relatively weak constraints on smoothness and joint safety, and without introducing noise. In the second stage, we refine the policies by tracking the motion trajectories obtained in the first stage, aiming for smoother transitions. We model the tracking problem as a bi-level optimization, where the tracking precision is dynamically adjusted based on the current tracking error, forming an adaptive curriculum mechanism. We apply this framework to a 1.7 m adult-scale humanoid robot, achieving stable execution in two representative real-world scenarios: sitting down onto a chair, stand up from a chair. Our approach provides a new perspective for the practical deployment of humanoid robots in real-world scenarios.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:

