On the convergence of projective-simulation-based reinforcement learning in Markov decision processes

W L Boyajian1, J Clausen1, L M Trenkwalder1

  • 1Institute for Theoretical Physics, University of Innsbruck, 6020 Innsbruck, Austria.

Quantum Machine Intelligence
|November 13, 2020
PubMed
Summary

Projective simulation, a quantum-inspired reinforcement learning approach, is formally analyzed. This study proves its convergence to optimal behavior in Markov decision processes, validating its theoretical performance.

Related Concept Videos

Decision Making: P-value Method01:09

Decision Making: P-value Method

The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.5K
Decision Making: Traditional Method01:14

Decision Making: Traditional Method

The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
4.9K
Decision Making01:20

Decision Making

Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
488
Reversible and Irreversible Processes01:14

Reversible and Irreversible Processes

The thermodynamic processes can be classified into reversible and irreversible processes. The processes that can be restored to their initial state are called reversible processes. It is only possible if the process is in quasi-static equilibrium, i.e., it takes place in infinitesimally small steps, and the system remains at equilibrium However, these are ideal processes and do not occur naturally. An ideal system undergoing a reversible process is always in thermodynamic equilibrium within...
5.3K
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
283
Reinforcement01:23

Reinforcement

Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
643