Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Self-Discrepancy Theory02:45

Self-Discrepancy Theory

18.3K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.  
18.3K
Instinctive Drift01:05

Instinctive Drift

208
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
208
Behaviorism01:28

Behaviorism

2.3K
The field of behaviorism was pioneered by figures such as Ivan Pavlov, John B. Watson, and B.F. Skinner fundamentally shifted the focus of psychology to the observable and controllable aspects of human and animal behavior. This shift marked a critical evolution in the discipline, emphasizing scientific rigor and experimental methodology.
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...
2.3K
Stereotype Threat and Self-fulfilling Prophecies02:09

Stereotype Threat and Self-fulfilling Prophecies

37.6K
When we hold a stereotype about a person, we have expectations that he or she will fulfill that stereotype. A self-fulfilling prophecy is an expectation held by a person that alters his or her behavior in a way that tends to make it true. When we hold stereotypes about a person, we tend to treat the person according to our expectations. This treatment can influence the person to act according to our stereotypic expectations, thus confirming our stereotypic beliefs. Research by Rosenthal and...
37.6K
Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

90
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
90
Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

532
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
532

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Progressive Epiglottic Deformity and Acute Supraglottic Edema: A Fatal Combination in Behçet's Disease.

Ear, nose, & throat journal·2026
Same author

Quantification of lettuce leaf DUS test traits and phenotypic fingerprint construction for variety identification.

Plant phenomics (Washington, D.C.)·2026
Same author

Dual-Modulus Microcone Array for Graded Tactile Sensing and Intelligent Slip Detection.

ACS applied materials & interfaces·2026
Same author

Deep learning-based 3D morphological segmentation and quantitative growth analysis of field-grown cabbage across the full cycle.

Plant phenomics (Washington, D.C.)·2026
Same author

A machine learning-helped antifouling strategy for improving the accuracy of electrochemical sensors.

Talanta·2026
Same author

Dual-Scale Synergistic Design: Oriented Material Stiffness and Deposition Path Planning for Enhanced Performance in Large-Format Additive Manufacturing of Short Carbon Fiber Components.

Materials (Basel, Switzerland)·2026

Related Experiment Video

Updated: Jun 26, 2025

A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
08:05

A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers

Published on: January 5, 2018

9.8K

Eliminating Primacy Bias in Online Reinforcement Learning by Self-Distillation.

Jingchen Li, Haobin Shi, Huarui Wu

    IEEE Transactions on Neural Networks and Learning Systems
    |May 17, 2024
    PubMed
    Summary

    Primacy bias in online reinforcement learning causes overfitting. Self-distillation reinforcement learning (SDRL) transfers knowledge to new policies, mitigating bias and improving generalization for better agent performance.

    More Related Videos

    Experimental Paradigm for Measuring the Effects of Self-distancing in Young Children
    07:01

    Experimental Paradigm for Measuring the Effects of Self-distancing in Young Children

    Published on: March 1, 2019

    7.9K
    Pavlovian Conditioned Approach Training in Rats
    06:57

    Pavlovian Conditioned Approach Training in Rats

    Published on: February 4, 2016

    11.0K

    Related Experiment Videos

    Last Updated: Jun 26, 2025

    A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
    08:05

    A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers

    Published on: January 5, 2018

    9.8K
    Experimental Paradigm for Measuring the Effects of Self-distancing in Young Children
    07:01

    Experimental Paradigm for Measuring the Effects of Self-distancing in Young Children

    Published on: March 1, 2019

    7.9K
    Pavlovian Conditioned Approach Training in Rats
    06:57

    Pavlovian Conditioned Approach Training in Rats

    Published on: February 4, 2016

    11.0K

    Area of Science:

    • Artificial Intelligence
    • Machine Learning

    Background:

    • Deep reinforcement learning is prone to overfitting early in training.
    • This overfitting, known as primacy bias, leads to suboptimal agent decisions and exploration.
    • Primacy bias obstructs agents from learning generalized policies for future states.

    Purpose of the Study:

    • To systematically investigate the causes and characteristics of primacy bias in online reinforcement learning.
    • To develop a novel framework, self-distillation reinforcement learning (SDRL), to overcome primacy bias.
    • To enhance policy generalization and accelerate learning in reinforcement learning agents.

    Main Methods:

    • Developed a self-distillation reinforcement learning (SDRL) framework based on knowledge distillation.
    • Agents transfer learned knowledge to a new, randomly initialized policy at regular intervals.
    • Incorporated L2 regularization and a self-imitation mechanism to prevent overfitting and accelerate learning.

    Main Results:

    • Experimental results in DMC and Atari 100k demonstrate SDRL's effectiveness in eliminating primacy bias.
    • The proposed method enables the generation of more generalized policies.
    • Policies refined through knowledge distillation lead to faster score improvements for agents.

    Conclusions:

    • Self-distillation reinforcement learning effectively addresses primacy bias in online reinforcement learning.
    • SDRL promotes the development of robust and generalized policies.
    • The framework enhances learning efficiency and agent performance in complex environments.