Related Experiment Video
Updated: Feb 2, 2026

Skeletal Phenotype Analysis of a Conditional Stat3 Deletion Mouse Model
Published on: July 3, 2020
The Boundedness Conditions for Model-Free HDP( λ )
Abstract:
This paper provides the stability analysis for a model-free action-dependent heuristic dynamic programing (HDP) approach with an eligibility trace long-term prediction parameter ( λ ). HDP( λ ) learns from more than one future reward. Eligibility traces have long been popular in Q-learning. This paper proves and demonstrates that they are worthwhile to use with HDP. In this paper, we prove its uniformly ultimately bounded (UUB) property under certain conditions. Previous works present a UUB proof for traditional HDP [HDP( λ = 0 )], but we extend the proof with the λ parameter. By using Lyapunov stability, we demonstrate the boundedness of the estimated error for the critic and actor neural networks as well as learning rate parameters. Three case studies demonstrate the effectiveness of HDP( λ ). The trajectories of the internal reinforcement signal nonlinear system are considered as the first case. We compare the results with the performance of HDP and traditional temporal difference [TD( λ )] with different λ values. The second case study is a single-link inverted pendulum. We investigate the performance of the inverted pendulum by comparing HDP( λ ) with regular HDP, with different levels of noise. The third case study is a 3-D maze navigation benchmark, which is compared with state action reward state action, Q( λ ), HDP, and HDP( λ ). All these simulation results illustrate that HDP( λ ) has a competitive performance; thus this contribution is not only UUB but also useful in comparison with traditional HDP.
More Related Videos
Related Concept Videos
Conditions on Early Earth
Classical Conditioning
Ivan Pavlov observed that dogs...
Conditions of Equilibrium
Internal forces are not considered for conditions of equilibrium because they occur in equal and opposite pairs within the body, effectively canceling each other. As a result,...
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Conditioned Taste Aversion
A notable characteristic of conditioned taste aversion is that it often requires only a single...
Electrostatic Boundary Conditions
The surface integral of an electric field is given by Gauss's law in integral form and is related to...

