Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reinforcement01:23

Reinforcement

192
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
192
Reinforcement Schedules01:24

Reinforcement Schedules

138
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
138
Observational Learning01:12

Observational Learning

155
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
155
Decision Making: P-value Method01:09

Decision Making: P-value Method

5.3K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.3K
Associative Learning01:27

Associative Learning

322
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
322
Masking and Demasking Agents01:19

Masking and Demasking Agents

2.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Design and Control of a 1-DOF MRI Compatible Pneumatically Actuated Robot with Long Transmission Lines.

IEEE/ASME transactions on mechatronics : a joint publication of the IEEE Industrial Electronics Society and the ASME Dynamic Systems and Control Division·2011
Same author

Effect of oxidized low-density lipoprotein concentration polarization on human smooth muscle cells' proliferation, cycle, apoptosis and oxidized low-density lipoprotein uptake.

Journal of the Royal Society, Interface·2011
Same author

Acrolein hydrogenation on Pt(211) and Au(211) surfaces: a density functional theory study.

Physical chemistry chemical physics : PCCP·2011
Same author

Anhydrous proton-conducting membrane based on poly-2-vinylpyridinium dihydrogenphosphate for electrochemical applications.

The journal of physical chemistry. B·2011
Same author

Pharmacophore identification, virtual screening and biological evaluation of prenylated flavonoids derivatives as PKB/Akt1 inhibitors.

European journal of medicinal chemistry·2011
Same author

Metabolomic study of insomnia and intervention effects of Suanzaoren decoction using ultra-performance liquid-chromatography/electrospray-ionization synapt high-definition mass spectrometry.

Journal of pharmaceutical and biomedical analysis·2011

Related Experiment Video

Updated: Jun 17, 2025

A Modified Lean and Release Technique to Emphasize Response Inhibition and Action Selection in Reactive Balance
07:19

A Modified Lean and Release Technique to Emphasize Response Inhibition and Action Selection in Reactive Balance

Published on: March 19, 2020

5.9K

Boosting Weak-to-Strong Agents in Multiagent Reinforcement Learning via Balanced PPO.

Sili Huang, Hechang Chen, Haiyin Piao

    IEEE Transactions on Neural Networks and Learning Systems
    |August 14, 2024
    PubMed
    Summary

    This study introduces dynamic policy balance (DPB) and weighted entropy regularization (WER) to address imbalanced training in multiagent reinforcement learning (RL). These methods improve individual policy learning and exploration efficiency for better overall performance.

    More Related Videos

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
    03:14

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

    Published on: December 6, 2024

    527
    Pavlovian Conditioned Approach Training in Rats
    06:57

    Pavlovian Conditioned Approach Training in Rats

    Published on: February 4, 2016

    10.9K

    Related Experiment Videos

    Last Updated: Jun 17, 2025

    A Modified Lean and Release Technique to Emphasize Response Inhibition and Action Selection in Reactive Balance
    07:19

    A Modified Lean and Release Technique to Emphasize Response Inhibition and Action Selection in Reactive Balance

    Published on: March 19, 2020

    5.9K
    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
    03:14

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

    Published on: December 6, 2024

    527
    Pavlovian Conditioned Approach Training in Rats
    06:57

    Pavlovian Conditioned Approach Training in Rats

    Published on: February 4, 2016

    10.9K

    Area of Science:

    • Artificial Intelligence
    • Machine Learning
    • Reinforcement Learning

    Background:

    • Multiagent policy gradients (MAPGs) are crucial in reinforcement learning (RL) but suffer from imbalanced individual policy training.
    • This imbalance, termed imbalance between policies (IBP), hinders overall multiagent system performance.

    Purpose of the Study:

    • To address the issue of imbalanced training in multiagent reinforcement learning.
    • To propose novel methods for balancing individual policy learning and enhancing exploration efficiency.

    Main Methods:

    • Proposed a dynamic policy balance (DPB) model that reweights training samples to balance individual policy learning.
    • Introduced weighted entropy regularization (WER) for team-level exploration with incentives for high-performing individuals.

    Main Results:

    • DPB and WER effectively alleviated imbalanced training across homogeneous and heterogeneous tasks.
    • The proposed methods demonstrated improved exploration efficiency compared to existing approaches.
    • Achieved an average performance gain of over 12.1% compared to state-of-the-art MAPG methods.

    Conclusions:

    • The proposed DPB and WER models offer a significant advancement in multiagent reinforcement learning.
    • These methods successfully address the critical challenge of imbalanced policy training and inefficient exploration.