Jove
Visualize
Contact Us

Related Concept Videos

Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
Observational Learning01:12

Observational Learning

Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Associative Learning01:27

Associative Learning

Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Avoidance Learning and Learned Helplessness01:14

Avoidance Learning and Learned Helplessness

Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Introduction to Learning01:18

Introduction to Learning

Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model01:13

Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model

Drugs administered through various routes can lead to nonlinear elimination, resulting in complex pharmacokinetic behaviors crucial to understanding efficacious drug dosing.
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Parallel algorithms for modules of learning automata.

IEEE transactions on systems, man, and cybernetics. Part B, Cybernetics : a publication of the IEEE Systems, Man, and Cybernetics Society·2008
Same author

Global Boltzmann perceptron network for online learning of conditional distributions.

IEEE transactions on neural networks·2008
Same author

Voronoi networks and their probability of misclassification.

IEEE transactions on neural networks·2008
Same author

Varieties of learning automata: an overview.

IEEE transactions on systems, man, and cybernetics. Part B, Cybernetics : a publication of the IEEE Systems, Man, and Cybernetics Society·2008
Same author

Analysis of the back-propagation algorithm with momentum.

IEEE transactions on neural networks·1994
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Learning the global maximum with parameterized learning automata.

M L Thathachar1, V V Phansalkar

  • 1Dept. of Electr. Eng., Indian Inst. of Sci., Bangalore.

IEEE Transactions on Neural Networks
|January 1, 1995
PubMed
Summary

This study introduces a decentralized reinforcement learning system using parameterized learning automata. The novel algorithm converges globally, optimizing functions effectively in simulations.

Related Experiment Videos

Area of Science:

  • Artificial Intelligence
  • Machine Learning
  • Computational Neuroscience

Background:

  • Reinforcement learning systems are crucial for artificial intelligence.
  • Decentralized algorithms offer advantages in scalability and robustness.
  • Learning automata provide a framework for adaptive decision-making.

Purpose of the Study:

  • To model a reinforcement learning system using a feedforward network of parameterized learning automata.
  • To introduce and analyze a novel decentralized update algorithm for learning automata.
  • To demonstrate the global convergence properties and practical performance of the proposed algorithm.

Main Methods:

  • A feedforward network architecture composed of teams of parameterized learning automata.
  • A decentralized update algorithm incorporating gradient-following and random perturbation terms.
  • Theoretical analysis of weak convergence to a Langevin equation solution.
  • Simulations on payoff games and pattern recognition tasks.

Main Results:

  • The proposed algorithm weakly converges to a solution of the Langevin equation.
  • The algorithm demonstrates global maximization of an appropriate objective function.
  • Decentralized units showed no information exchange during updates.
  • Reasonable convergence rates were achieved in simulations.

Conclusions:

  • The developed algorithm provides a robust and effective method for decentralized reinforcement learning.
  • The model successfully links learning automata dynamics to stochastic differential equations.
  • The approach shows promise for applications in game theory and pattern recognition.