Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant factor...
Reinforcement Schedules01:24

Reinforcement Schedules

Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Observational Learning01:12

Observational Learning

Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Test of CP Symmetry in the Neutral Decays of Λ via J/ψ→ΛΛ[over ¯].

Physical review letters·2026
Same author

Precise Measurement of the Chromoelectric Dipole Moment of the Charm Quark.

Physical review letters·2026
Same author

Precise Measurement of Matter-Antimatter Asymmetry with Entangled Hyperon-Antihyperon Pairs.

Physical review letters·2026
Same author

[Analysis of 3-year postoperative prognosis and influencing factors in patients with H3 K27-altered diffuse midline glioma].

Zhonghua yi xue za zhi·2026
Same author

Observation of Λ[over ¯]p→K^{+}π^{+}π^{-}π^{0} and Λ[over ¯]p→K^{+}π^{+}π^{-}2π^{0}.

Physical review letters·2026
Same author

First Measurement of the D_{s}^{+}→K^{0}μ^{+}ν_{μ} Decay.

Physical review letters·2026

Related Experiment Videos

Integrating temporal difference methods and self-organizing neural networks for reinforcement learning with delayed

A H Tan1, N Lu, D Xiao

  • 1School of Computer Engineering and Intelligent Systems Centre, Nanyang Technological University, Singapore. asahtan@ntu.edu.sg

IEEE Transactions on Neural Networks
|February 14, 2008
PubMed
Summary

This study introduces TD-FALCON, a neural network model that integrates adaptive resonance theory and temporal difference learning for autonomous agents. TD-FALCON efficiently learns optimal navigation strategies in dynamic environments using reinforcement signals.

Related Experiment Videos

Area of Science:

  • Artificial Intelligence
  • Computational Neuroscience
  • Robotics

Background:

  • Autonomous agents require robust learning mechanisms for dynamic environments.
  • Integrating multimodal sensory inputs, actions, and rewards is crucial for effective decision-making.
  • Existing reinforcement learning systems often struggle with delayed feedback and computational efficiency.

Purpose of the Study:

  • To present a novel neural architecture, TD-FALCON, for learning category node encodings across multimodal patterns.
  • To enable autonomous agents to adapt and function effectively using immediate and delayed reinforcement signals.
  • To enhance learning speed, task completion, and efficiency in agent navigation.

Main Methods:

  • Developed the TD fusion architecture for learning, cognition, and navigation (TD-FALCON).
  • Integrated adaptive resonance theory (ART) and temporal difference (TD) learning methods.
  • Utilized on-policy (SARSA) and off-policy (Q-learning) TD methods to estimate value functions.
  • Implemented an action selection policy based on learned value functions.

Main Results:

  • TD-FALCON systems demonstrated effective learning with both immediate and delayed reinforcement.
  • Achieved stable performance significantly faster than standard gradient-descent-based reinforcement learning systems.
  • Exhibited superior task completion rates and efficiency in a minefield navigation task.

Conclusions:

  • TD-FALCON provides an efficient neural architecture for reinforcement learning in dynamic environments.
  • The integration of ART and TD learning enhances an agent's adaptability and learning speed.
  • TD-FALCON shows promise for developing more capable and efficient autonomous agents.