Related Experiment Video
Updated: Nov 19, 2025

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
Published on: May 8, 2021
Examining the Use of Temporal-Difference Incremental Delta-Bar-Delta for Real-World Predictive Knowledge
Johannes Günther1,2, Nadia M Ady1, Alex Kearney1
1Department of Computing Science, University of Alberta, Edmonton, AB, Canada.
This study introduces Temporal-Difference Incremental Delta-Bar-Delta (TIDBD) for robot learning, enabling adaptive learning rates and improved prediction accuracy. TIDBD offers a robust alternative to traditional methods, even detecting sensor failures in robotic systems.
Area of Science:
- Robotics
- Machine Learning
- Control Systems
Background:
- Predictive knowledge enhances robot control and other applications.
- Online, incremental learning through environmental interaction is key for robotics.
- A challenge is selecting appropriate learning parameters, such as learning rates or step sizes.
Purpose of the Study:
- To examine online step-size adaptation using Temporal-Difference Incremental Delta-Bar-Delta (TIDBD).
- To evaluate TIDBD as a practical alternative to classic Temporal-Difference (TD) learning.
- To demonstrate TIDBD's ability to detect sensor failures in robotic applications.
Main Methods:
- Applied TIDBD to a Modular Prosthetic Limb, a sensor-rich robotic arm.
- TIDBD learns and adapts step sizes on a feature level, enabling simultaneous step-size tuning and representation learning.
- Compared TIDBD's performance to classic TD learning with extensive parameter searches.
Main Results:
- TIDBD performs comparably to hand-tuned TD learning in predicting robotic data streams.
- TIDBD automatically detects patterns indicative of sensor failures, a common issue in robotic applications.
- TIDBD demonstrates robustness to initial step-size values, outperforming classic TD.
Conclusions:
- TIDBD is a practical and robust method for online step-size adaptation in robotic learning.
- The approach improves the ability of robotic devices to learn from environmental interactions.
- These findings enhance capabilities for autonomous agents and robots.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Associative Learning
Classical conditioning, also known...
Observational Learning
PD Controller: Design
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Multi-input and Multi-variable systems
In the absence of...
Time-Domain Interpretation of PD Control
Consider the example of control of motor torque. Initially, a positive...
