Related Experiment Video
Updated: Jan 26, 2026

11:18
Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task
Published on: June 1, 2015
11.1K
The Hierarchical Continuous Pursuit Learning Automation: A Novel Scheme for Environments With Large Numbers of
IEEE Transactions on Neural Networks and Learning Systems
|April 17, 2019
Summary
This study introduces a hierarchical learning automata (LA) solution for problems with many actions, overcoming limitations of traditional methods. The novel approach ensures convergence to optimal actions, even with vast action spaces.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Reinforcement Learning
Background:
- Traditional learning automata (LA) methods struggle with environments featuring a large number of actions due to high-dimensional action probability vectors.
- Very large action spaces lead to action probabilities becoming smaller than machine accuracy, preventing their selection and hindering effective learning.
Purpose of the Study:
- To present a novel hierarchical solution extending the continuous pursuit paradigm to learning automata problems with a large number of actions.
- To address the challenge of arbitrarily small action probabilities in high-dimensional action spaces.
Main Methods:
- A hierarchical structure is proposed where all environment actions are leaves.
- Each level of the hierarchy utilizes a two-action learning automaton.
- The pursuit paradigm is employed at each level, allowing the best action to propagate upwards towards the root.
Main Results:
- Formal proof of ϵ-optimal convergence for the hierarchical scheme.
- Extensive experimental validation on environments with up to 256 actions.
- Demonstrated computational advantages, with the hierarchical continuous pursuit automaton requiring up to 82% fewer iterations than the benchmark LR-I scheme.
Conclusions:
- The proposed hierarchical learning automata scheme effectively solves problems with large action spaces.
- The method offers significant computational advantages and converges to optimal actions.
- The approach is extendable to discretized and Bayesian pursuit learning automata.
Related Concept Videos
The Z-Scheme of Electron Transport in Photosynthesis
13.4K
The light reactions of photosynthesis assume a linear flow of electrons from water to NADP+. During this process, light energy drives the splitting of water molecules to produce oxygen. However, oxidation of water molecules is a thermodynamically unfavorable reaction and requires a strong oxidizing agent. This is accomplished by the first product of light reactions: oxidized P680 (or P680+), the most powerful oxidizing agent known in biology. The oxidized P680 that acquires an electron from the...
13.4K
Fixed Action Patterns
17.6K
A fixed action pattern (FAP) is a specific, hard-wired sequence of behaviors that occurs in response to an external stimulus, called a sign stimulus. The behavior is “fixed” because it is essentially unchangeable—proceeding similarly across individuals of a species every time it occurs.
17.6K
Action Potential
4.4K
Neurons communicate by firing action potentials—the electrochemical signal that is propagated along the axon. The signal results in the release of neurotransmitters at axon terminals, thereby transmitting information to the nervous system. An action potential is a specific "all-or-none" change in membrane potential that results in a rapid spike in voltage.
Membrane potential in neurons
Neurons typically have a resting membrane potential of about -70 millivolts (mV). When they receive...
Membrane potential in neurons
Neurons typically have a resting membrane potential of about -70 millivolts (mV). When they receive...
4.4K
Action Potential
10.9K
Neurons communicate by firing action potentials—the electrochemical signal that is propagated along the axon. The signal results in the release of neurotransmitters at axon terminals, thereby transmitting information to the nervous system. An action potential is a specific "all-or-none" change in membrane potential that results in a rapid spike in voltage.
Membrane potential in neurons
Neurons typically have a resting membrane potential of about -70 millivolts (mV). When they receive...
Membrane potential in neurons
Neurons typically have a resting membrane potential of about -70 millivolts (mV). When they receive...
10.9K
Action Potentials
141.8K
Overview
141.8K
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K

