Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Decision Making: P-value Method01:09

Decision Making: P-value Method

7.1K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
7.1K
Reinforcement01:23

Reinforcement

1.1K
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
1.1K
Approximate Integration01:24

Approximate Integration

92
In many practical and theoretical contexts, the exact value of a definite integral may be inaccessible. This limitation typically arises when the antiderivative of a function is either unknown or cannot be expressed in a closed mathematical form. Alternatively, it can occur when a function is defined not by a formula but by a finite set of empirical data points, such as those collected during experiments. In these cases, approximate integration techniques provide a valuable solution.One of the...
92
Reinforcement Schedules01:24

Reinforcement Schedules

664
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
664
Linearization and Approximation01:26

Linearization and Approximation

130
Linearization is a mathematical technique used to approximate complex, nonlinear functions with simpler linear models in the vicinity of a chosen reference point. The method is based on the idea that, although a function may be difficult to evaluate exactly, its behavior near a specific input value can often be closely approximated by the tangent line at that point. This approach is particularly useful when small deviations from a known value are involved.Consider the square root function, for...
130
Expected Value01:15

Expected Value

8.1K
The expected value is known as the "long-term" average or mean. This means that over the long term of experimenting over and over, you would expect this average. The expected average is represented by the symbol μ. It is calculated as follows:
8.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Circular inference in visual cortical hierarchies in schizophrenia.

Psychiatry and clinical neurosciences·2026
Same author

Data-driven inverse optimal control for continuous-time nonlinear systems.

ISA transactions·2025
Same author

A computational model of canonical cortical microcircuits for dynamic Bayesian inference and control as inference.

Neuroscience research·2025
Same author

Possible contribution to data-driven primate research: Comment on "Kinematic coding: Measuring information in naturalistic behaviour" by Becchio, Pullar, Scaliti, and Panzeri.

Physics of life reviews·2025
Same author

Optical Neuroimage Studio (OptiNiSt): Intuitive, scalable, extendable framework for optical neuroimage data analysis.

PLoS computational biology·2025
Same author

Information-Theoretical Analysis of Team Dynamics in Football Matches.

Entropy (Basel, Switzerland)·2025

Related Experiment Videos

From free energy to expected energy: Improving energy-based value function approximation in reinforcement learning.

Stefan Elfwing1, Eiji Uchibe1, Kenji Doya2

  • 1Department of Brain Robot Interface, ATR Computational Neuroscience Laboratories, 2-2-2 Hikaridai, Seikacho, Soraku-gun, Kyoto 619-0288, Japan; Okinawa Institute of Science and Technology Graduate University, 1919-1 Tancha, Onna-son, Okinawa 904-0495, Japan.

Neural Networks : the Official Journal of the International Neural Network Society
|September 19, 2016
PubMed
Summary

Expected Energy Reinforcement Learning (EERL) improves upon Free-Energy based Reinforcement Learning (FERL) for complex tasks. EERL handles continuous inputs and achieves superior performance in various challenging reinforcement learning benchmarks.

Keywords:
Expected energyFunction approximationReinforcement learningRestricted Boltzmann machineSZ-Tetris

Related Experiment Videos

Area of Science:

  • Artificial Intelligence
  • Machine Learning
  • Reinforcement Learning

Background:

  • Free-Energy based Reinforcement Learning (FERL) is effective for high-dimensional spaces but limited to binary or near-binary state inputs.
  • FERL approximates value functions using the negative free energy of Restricted Boltzmann Machines (RBMs).
  • Previous work showed FERL robustness improves with free energy scaling related to network size.

Purpose of the Study:

  • To enhance Restricted Boltzmann Machine (RBM) function approximation in reinforcement learning.
  • To introduce a novel method, Expected Energy Reinforcement Learning (EERL), capable of handling continuous state inputs.
  • To improve upon the performance and applicability of Free-Energy based Reinforcement Learning (FERL).

Main Methods:

  • Proposed approximating the value function with the negative expected energy (EERL) instead of negative free energy.
  • Utilized Restricted Boltzmann Machines (RBMs) for function approximation.
  • Validated EERL across gridworld tasks, SZ-Tetris, and a robot navigation task with image-based inputs.

Main Results:

  • EERL outperformed FERL, standard neural networks, and linear function approximation on high-dimensional gridworld tasks.
  • EERL achieved state-of-the-art results in both model-free and model-based learning for stochastic SZ-Tetris.
  • EERL significantly surpassed FERL and standard neural network approximation in a robot navigation task using raw, noisy RGB images.

Conclusions:

  • Expected Energy Reinforcement Learning (EERL) offers a more robust and versatile approach compared to FERL.
  • EERL effectively handles continuous state inputs, expanding the applicability of energy-based models in reinforcement learning.
  • The proposed EERL method demonstrates significant performance improvements across diverse and challenging reinforcement learning problems.