Related Experiment Videos
Recovering Reward Functions From Distributed Expert Demonstrations via Bi-Level Maximum-Likelihood Optimization
Summary
Federated maximum-likelihood IRL (F-ML-IRL) enables decentralized reward inference from expert data. This novel algorithm ensures convergence and outperforms centralized methods in robotic control tasks.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Inverse reinforcement learning (IRL) infers reward functions and policies from expert demonstrations.
- Current IRL methods often require centralized data access, posing challenges for decentralized and privacy-sensitive applications.
Purpose of the Study:
- To propose a novel federated maximum-likelihood IRL (F-ML-IRL) algorithm for decentralized reward inference.
- To analyze the convergence rate of the proposed F-ML-IRL algorithm.
Main Methods:
- F-ML-IRL utilizes dual aggregation for global model updates.
- Bi-level local updates optimize reward functions and agent policies using maximum likelihood and entropy regularization.
Main Results:
- The F-ML-IRL algorithm's global model converges to a stationary point for reward and policy parameters in finite time.
- Demonstrated convergence of recovered rewards in decentralized learning settings.
- Outperformed centralized baselines in 12 out of 20 high-dimensional robotic control tasks.
Conclusions:
- F-ML-IRL effectively addresses the limitations of centralized IRL in decentralized environments.
- The algorithm ensures convergence and achieves superior performance by leveraging distributed data.