Related Experiment Video
Updated: Oct 7, 2025

07:52
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
8.9K
A Critical Period for Robust Curriculum-Based Deep Reinforcement Learning of Sequential Action in a Robot Arm.
Roy de Kleijn1, Deniz Sen2, George Kachergis3
1Leiden Institute for Brain and Cognition, Leiden University.
Topics in Cognitive Science
|January 10, 2022
Summary
Virtual agents trained with a curriculum, starting with rewards and then adding energy penalties, developed human-like centering behavior in sequential actions. Early exploration is key for robust learning, similar to infant curiosity.
Area of Science:
- Robotics
- Cognitive Science
- Machine Learning
Background:
- Everyday activities often involve sequential actions with context effects influencing execution.
- Centering behavior, a strategy to minimize movement time by positioning equidistant to targets, is observed in human sequential action tasks.
- Investigating sequential action learning in artificial agents can provide insights into human motor control and learning.
Purpose of the Study:
- To investigate if virtual robotic agents can develop human-like sequential action strategies, specifically centering behavior.
- To explore the impact of different training methodologies, including curricularized learning and energy expenditure penalties, on agent performance.
- To understand the role of early exploration and potential critical periods in the development of sequential action learning in artificial agents.
Main Methods:
- Trained a virtual robotic agent using proximal policy optimization (a deep reinforcement learning algorithm).
- Designed a task mimicking the human serial response time task, requiring agents to reach for appearing targets.
- Implemented a curriculum: initially rewarding target reaching, then introducing an energy expenditure penalty.
Main Results:
- Agents trained with the curriculum were more likely to develop centering behavior, mirroring human strategies.
- Agents trained without a curriculum (penalty introduced immediately) showed poor learning and high performance variability due to limited exploration.
- Early energetic exploration, facilitated by the curriculum, promoted more robust learning in agents.
Conclusions:
- Curricularized learning, emphasizing early exploration, enables virtual agents to develop sophisticated sequential action strategies like centering behavior.
- Findings suggest parallels between infant learning (curiosity-driven exploration) and optimal training strategies for artificial agents.
- The study highlights the importance of developmental timing in learning, indicating potential critical periods for incorporating new objectives in both humans and agents.
Related Concept Videos
Observational Learning
362
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
362
Reinforcement Schedules
257
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
257
Muscle Coordination and Action
2.2K
Muscle coordination is a complex and finely tuned process essential for smooth and purposeful movements like flexion, extension, adduction, abduction, and rotation. The human body orchestrates the actions of various muscles working in concert, each with a specific role. Four functional types describe how muscles work together: agonist, antagonist, synergist, and fixator.
Agonists
Agonist muscles, often called prime movers, are the primary muscles responsible for producing a specific movement....
Agonists
Agonist muscles, often called prime movers, are the primary muscles responsible for producing a specific movement....
2.2K
Reinforcement
420
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
420

