Related Experiment Video
Updated: Jun 15, 2025

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
Published on: May 8, 2021
Active Inference and Reinforcement Learning: A Unified Inference on Continuous State and Action Spaces Under Partial
Parvin Malekzadeh1, Konstantinos N Plataniotis2
1Edward S. Rogers Sr. Department of Electrical and Computer Engineering, University of Toronto, M5S 3G8, Canada p.malekzadeh@mail.utoronto.ca.
This study unifies reinforcement learning (RL) and active inference (AIF) to create better decision-making agents for partially observable environments. The new approach improves learning in continuous spaces and makes reward design optional.
Area of Science:
- Artificial Intelligence
- Computational Neuroscience
Background:
- Reinforcement learning (RL) excels in fully observable environments but struggles with partial observations common in real-world scenarios.
- Partially Observable Markov Decision Processes (POMDPs) model these complex environments, but traditional RL methods face challenges with long time horizons and high-dimensional data.
- Active Inference (AIF) offers an alternative by minimizing expected free energy (EFE), balancing exploration and exploitation, yet is limited by computational demands in large spaces.
Purpose of the Study:
- To propose a unified principle connecting RL and AIF for enhanced agent decision-making in POMDPs.
- To overcome the limitations of existing RL and AIF methods in continuous, high-dimensional, and long-horizon partially observable environments.
- To demonstrate a novel approach that integrates reward-maximizing and information-seeking behaviors.
Main Methods:
- Developed a theoretical framework unifying RL and AIF principles.
- Formulated a novel approach applicable to continuous space POMDPs.
- Conducted rigorous theoretical analysis and experimental validation.
Main Results:
- Established a theoretical connection between AIF and RL, enabling seamless integration.
- Demonstrated superior learning capabilities in continuous space POMDPs compared to existing RL methods.
- Showcased the ability to solve reward-free problems by leveraging information-seeking exploration.
Conclusions:
- The unified principle effectively bridges RL and AIF, overcoming limitations of individual approaches.
- The proposed method offers a powerful new tool for designing artificial agents capable of handling complex, partially observable environments.
- This work opens new avenues for AIF in artificial intelligence, reducing reliance on explicit reward specification.
Related Concept Videos
Observational Learning
State Space Representation
Consider an RLC circuit, a...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Associative Learning
Classical conditioning, also known...
Reinforcement Schedules
Once a behavior is learned,...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...

