Related Experiment Video
Updated: May 11, 2026

11:18
Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task
Published on: June 1, 2015
Humanoids Learning to Walk: A Natural CPG-Actor-Critic Architecture.
Cai Li1, Robert Lowe, Tom Ziemke
1Interaction Lab, University of Skövde Skövde, Sweden.
Frontiers in Neurorobotics
|May 16, 2013
Summary
This study presents a new model for humanoid robots to learn walking using dynamic systems theory (DST) and reinforcement learning (RL). The approach integrates a central pattern generator (CPG) with adaptive value functions for effective environmental interaction.
Area of Science:
- Robotics
- Computational Neuroscience
- Artificial Intelligence
Background:
- Humanoid locomotion learning is challenging.
- Dynamic Systems Theory (DST) offers a novel approach via environmental interaction.
- Reinforcement Learning (RL) provides adaptive linking through reward systems.
Purpose of the Study:
- Propose an integrated model combining DST and RL for humanoid robot walking.
- Apply the model to a NAO robot learning to walk.
- Demonstrate emergent walking ability from value-based environmental interaction.
Main Methods:
- Integrated a simplified central pattern generator (CPG) architecture with an actor-critic RL approach (cpg-actor-critic).
- Employed least-square-temporal-difference learning with natural gradient for fast convergence and balanced exploration/exploitation.
- Utilized a dynamic value function as an adaptive stability indicator instead of traditional rewards.
Main Results:
- The cpg-actor-critic model successfully enabled a NAO robot to learn walking.
- The emergent walking ability stemmed from value-based interaction with the environment.
- Analysis using a DST-based embodied cognition approach provided insights into the learning process.
Conclusions:
- Robot walking emerges from integrating sensorimotor activity and value.
- The proposed model effectively combines DST, RL, and neuroscientific principles.
- This approach offers a promising direction for humanoid robot learning and embodied cognition research.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
Purposive Learning
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a bonus...
Fixed Action Patterns
A fixed action pattern (FAP) is a specific, hard-wired sequence of behaviors that occurs in response to an external stimulus, called a sign stimulus. The behavior is “fixed” because it is essentially unchangeable—proceeding similarly across individuals of a species every time it occurs.
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Cognitive Learning
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Actor-Observer Effect
The actor-observer effect, a cognitive bias closely linked to the fundamental attribution error, refers to the tendency for individuals to attribute their behavior to external, situational factors while explaining others’ behavior in terms of internal, dispositional traits. This asymmetry in attribution significantly influences social perception and judgment.Cognitive Mechanisms Behind the EffectTwo primary psychological mechanisms contribute to the actor-observer effect: differences in visual...
