Related Experiment Video
Updated: Jan 16, 2026

07:52
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
9.2K
DT-HRL: Mastering Long-Sequence Manipulation with Reimagined Hierarchical Reinforcement Learning
Junyang Zhang1, Yilin Zhang1, Honglin Sun1
1Graduate School of Information, Production and Systems, Waseda University, Kitakyushu 808-0135, Japan.
Biomimetics (Basel, Switzerland)
|September 26, 2025
Summary
This study introduces a Hierarchical Reinforcement Learning (HRL) framework using a Decision Transformer (DT) for robotic manipulators. The new approach enhances long-term reasoning and generalization in complex logistics tasks.
Area of Science:
- Robotics
- Artificial Intelligence
- Machine Learning
Background:
- Robotic manipulators in logistics face challenges with multi-step tasks, frequent switching, and long-term dependencies.
- Existing methods struggle with complex, sequential decision-making in dynamic environments.
- Human motor control offers a model for hierarchical task execution.
Purpose of the Study:
- To propose a novel Hierarchical Reinforcement Learning (HRL) framework for robotic manipulators.
- To improve long-term reasoning, generalization, and task execution in logistics.
- To integrate Decision Transformer (DT) capabilities with hierarchical control structures.
Main Methods:
- A multi-task goal-conditioned Decision Transformer (MTGC-DT) framework was developed.
- A high-level policy models the Markov decision process as a sequence modeling task.
- A low-level policy utilizes parameterized action primitives for physical execution.
- Introduced a path-efficiency loss (PEL) correction and a learnable primitive skill library.
Main Results:
- The Decision Transformer-based Hierarchical Reinforcement Learning (DT-HRL) achieved over 10% higher success rate compared to baselines.
- DT-HRL demonstrated over 8% higher average reward in logistics tasks.
- Ablation experiments showed a normalized score increase of over 2%.
Conclusions:
- The proposed DT-HRL framework effectively addresses long-term dependencies and improves generalization in robotic manipulation.
- Integrating Decision Transformers with HRL offers a promising direction for complex task automation in logistics.
- The framework's modular design with parameterized skills enhances reusability and adaptability.
Related Concept Videos
Reinforcement Schedules
459
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
459
Observational Learning
838
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
838
Reinforcement
839
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
839
Elaborative Rehearsals
338
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
338
Associative Learning
1.2K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
1.2K
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
3.8K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.8K

