Related Experiment Video
Updated: May 21, 2026

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
Dynamics of Boltzmann Q learning in two-player two-action games
Ardeshir Kianercy1, Aram Galstyan
1USC Information Sciences Institute, Marina del Rey, California 90292, USA.
Summary
Q learning dynamics in two-player games with exploration do not always converge to Nash equilibria (NEs). Learning can lead to unique rest points, especially with multiple NEs or specific exploration rates.
Area of Science:
- Game theory
- Reinforcement learning
- Computational economics
Background:
- Q learning is a reinforcement learning algorithm used to find optimal strategies in games.
- Boltzmann exploration is a common mechanism for introducing stochasticity in agent decision-making.
- Nash equilibria (NEs) represent stable strategy profiles in game theory where no player can unilaterally improve their outcome.
Purpose of the Study:
- To analyze the convergence properties of Q learning in two-player, two-action games with Boltzmann exploration.
- To characterize the structure of rest points in Q learning dynamics and their relationship to Nash equilibria (NEs).
- To investigate the impact of exploration rates on the stability and existence of these rest points.
Main Methods:
- Mathematical analysis of Q learning dynamics under a Boltzmann exploration mechanism.
- Characterization of the system's rest points for various game structures.
- Sensitivity analysis of rest point structure with respect to exploration noise.
Main Results:
- For any non-zero exploration rate, Q learning dynamics are dissipative, ensuring convergence to rest points.
- These rest points often differ from the game's Nash equilibria (NEs).
- Drastic changes in learning dynamics occur at critical exploration rates for games with multiple NEs.
- Unique, non-NE rest points can emerge and persist within specific exploration rate ranges for games with a single NE.
Conclusions:
- Q learning with Boltzmann exploration does not guarantee convergence to Nash equilibria (NEs) in two-player games.
- The exploration mechanism significantly influences the convergence behavior and the nature of the resulting stable states.
- Understanding these dynamics is crucial for designing effective learning agents in strategic environments.
Related Concept Videos
Maxwell-Boltzmann Distribution: Problem Solving
Individual molecules in a gas move in random directions, but a gas containing numerous molecules has a predictable distribution of molecular speeds, which is known as the Maxwell-Boltzmann distribution, f(v).
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
This distribution function f(v) is defined by saying that the expected number N (v1,v2) of particles with speeds between v1 and v2 is given by
Reaction Quotient
The status of a reversible reaction is conveniently assessed by evaluating its reaction quotient (Q). For a reversible reaction described by m A + n B ⇌ x C + y D, the reaction quotient is derived directly from the stoichiometry of the balanced equation as
Second Order systems II
In an underdamped second-order system, where the damping ratio ζ is between 0 and 1, a unit-step input results in a transfer function that, when transformed using the inverse Laplace method, reveals the output response. The output exhibits a damped sinusoidal oscillation, and the difference between the input and output is termed the error signal. This error signal also demonstrates damped oscillatory behavior. Eventually, as the system reaches a steady state, the error diminishes to zero.
If ζ...
If ζ...
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
First Law: Particles in Two-dimensional Equilibrium
Recall that a particle in equilibrium is one for which the external forces are balanced. Static equilibrium involves objects at rest, and dynamic equilibrium involves objects in motion without acceleration; but it is important to remember that these conditions are relative. For instance, an object may be at rest when viewed from one frame of reference, but that same object would appear to be in motion when viewed by someone moving at a constant velocity.
Newton's first law tells us about the...
Newton's first law tells us about the...
Dynamic Equilibrium
A reversible chemical reaction represents a chemical process that proceeds in both forward (left to right) and reverse (right to left) directions. When the rates of the forward and reverse reactions are equal, the concentrations of the reactant and product species remain constant over time and the system is at equilibrium. A special double arrow is used to emphasize the reversible nature of the reaction. The relative concentrations of reactants and products in equilibrium systems vary greatly;...