Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reinforcement01:23

Reinforcement

199
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
199
Purposive Learning01:22

Purposive Learning

107
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
107
Observational Learning01:12

Observational Learning

158
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
158
Propagation of Action Potentials01:23

Propagation of Action Potentials

5.5K
The propagation of an action potential refers to the process by which a nerve impulse, or "action potential," travels along a neuron.
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
5.5K
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

105
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
105
Reinforcement Schedules01:24

Reinforcement Schedules

139
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
139

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Microgels prepared by microfluidics from structural design to practical applications: Development and challenge.

Advances in colloid and interface science·2026
Same author

Ultrasound-Assisted Covalent Conjugation of Walnut Albumin with Bound Polyphenols: Structural Modulation and Functional Enhancement.

Foods (Basel, Switzerland)·2026
Same author

Digital twin-centered food safety management systems: A review of IoT, AI, and blockchain integration for bacterial pathogen control.

Food research international (Ottawa, Ont.)·2026
Same author

Inactivating Bacillus cereus spores with pulsed light: roles of thermal and reactive oxygen species-mediated mechanisms.

Food research international (Ottawa, Ont.)·2026
Same author

Structural similarity analysis and AI-predicted binding mode of bitter flavonoids in pomelo (Citrus maxima) using HRMS-based metabolomics and molecular fingerprinting.

Food research international (Ottawa, Ont.)·2026
Same author

<i>Bifidobacterium animalis</i> ssp. <i>lactis</i> 420 and <i>Cordyceps militaris</i> Synergistically Modulate the Gut Microbiota by Increasing Mucin 2 Production.

Nutrients·2026

Related Experiment Video

Updated: Jun 18, 2025

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

8.7K

Generative subgoal oriented multi-agent reinforcement learning through potential field.

Shengze Li1, Hao Jiang1, Yuntao Liu1

  • 1Academy of Military Science, Beijing, 100000, China.

Neural Networks : the Official Journal of the International Neural Network Society
|August 1, 2024
PubMed
Summary

This study introduces a novel Potential field Subgoal-based Multi-Agent reinforcement learning (PSMA) method to unify learning objectives in multi-agent reinforcement learning (MARL). PSMA enhances agent learning speed and effectiveness in sparse reward tasks by using potential fields for subgoal generation and achievement.

Keywords:
Multi-agent reinforcement learningPotential fieldSubgoal generation

More Related Videos

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
11:54

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

Published on: May 8, 2021

4.3K
The HoneyComb Paradigm for Research on Collective Human Behavior
06:48

The HoneyComb Paradigm for Research on Collective Human Behavior

Published on: January 19, 2019

9.3K

Related Experiment Videos

Last Updated: Jun 18, 2025

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

8.7K
Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
11:54

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

Published on: May 8, 2021

4.3K
The HoneyComb Paradigm for Research on Collective Human Behavior
06:48

The HoneyComb Paradigm for Research on Collective Human Behavior

Published on: January 19, 2019

9.3K

Area of Science:

  • Artificial Intelligence
  • Machine Learning
  • Robotics

Background:

  • Multi-agent reinforcement learning (MARL) utilizes subgoals to accelerate agent learning in sparse reward environments.
  • Existing MARL methods often lack consistency between subgoal generation and achievement, hindering learning effectiveness.

Purpose of the Study:

  • To propose a novel Potential field Subgoal-based Multi-Agent reinforcement learning (PSMA) method.
  • To unify the learning objectives of subgoal generation and subgoal reached stages in MARL.

Main Methods:

  • Introduced potential fields (PF) to represent agent states and measure inter-agent interactions.
  • Developed a state-to-PF representation model and a subgoal selector using an experience replay buffer.
  • Defined an intrinsic reward function to guide agents toward subgoals while maximizing joint action-value.

Main Results:

  • The proposed PSMA method effectively unifies two-stage learning objectives.
  • PSMA demonstrated superior performance compared to state-of-the-art MARL methods.
  • Achieved significant improvements on StarCraft II micro-management (SMAC) and Google Research Football (GRF) tasks in sparse reward settings.

Conclusions:

  • The PSMA method offers a unified approach to subgoal learning in MARL.
  • Potential fields provide an effective mechanism for representing agent states and interactions.
  • PSMA significantly enhances MARL performance in complex, sparse reward environments.