Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

702
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
702
Law of Effect01:06

Law of Effect

1.5K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.5K
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model01:13

Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model

103
Drugs administered through various routes can lead to nonlinear elimination, resulting in complex pharmacokinetic behaviors crucial to understanding efficacious drug dosing.
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
103
Purposive Learning01:22

Purposive Learning

178
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
178
Role of Shaping in Operant Conditioning01:19

Role of Shaping in Operant Conditioning

439
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
439
Reinforcement01:23

Reinforcement

307
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
307

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Hydrostatic pressure spectroscopy reveals the adaptation of microbial rhodopsins to high-pressure environment.

Scientific reports·2026
Same author

Solution structure of mouse HBS1L/SKI7-specific UBA domain in complex with ubiquitin: Implications for stalled ribosome recognition.

PloS one·2026
Same author

NMR characterization of the structure and interaction of an RNA aptamer targeting α-synuclein.

Biochemical and biophysical research communications·2026
Same author

Structural insights into the interaction between the BH3-like domain of hepatitis B virus X protein and LC3B.

Biochimica et biophysica acta. Proteins and proteomics·2026
Same author

Impact of Nucleotide Flexibility on Aptamer-Protein Recognition: RNA vs RNA-DNA Chimera.

ACS chemical biology·2026
Same author

Contrasting effects of different molecular crowding environments on base-pair opening/closing dynamics of DNA triplex structures.

The FEBS journal·2026

Related Experiment Video

Updated: Aug 10, 2025

Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

11.0K

Achieving efficient interpretability of reinforcement learning via policy distillation and selective input gradient

Jinwei Xing1, Takashi Nagata2, Xinyun Zou2

  • 1Department of Cognitive Sciences, University of California, Irvine, 92697, CA, USA.

Neural Networks : the Official Journal of the International Neural Network Society
|February 12, 2023
PubMed
Summary

This study introduces Distillation with selective Input Gradient Regularization (DIGR) for interpretable deep Reinforcement Learning (RL). DIGR enhances policy interpretability and computational efficiency, while also improving robustness against adversarial attacks.

Keywords:
Efficient interpretabilityInterpretable reinforcement learningSaliency map

More Related Videos

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
11:54

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

Published on: May 8, 2021

4.5K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.3K

Related Experiment Videos

Last Updated: Aug 10, 2025

Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

11.0K
Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
11:54

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

Published on: May 8, 2021

4.5K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.3K

Area of Science:

  • Artificial Intelligence
  • Machine Learning
  • Deep Learning

Background:

  • Deep Reinforcement Learning (RL) excels in various tasks but lacks interpretability for real-world applications.
  • Saliency maps are common for deep neural network interpretability, but existing RL methods are computationally expensive or ineffective.
  • Current saliency map techniques for RL struggle with real-time performance and policy interpretability.

Purpose of the Study:

  • To develop an interpretable and computationally efficient method for generating saliency maps in deep RL.
  • To enhance the robustness of RL policies against adversarial attacks.
  • To validate the proposed approach across diverse RL tasks.

Main Methods:

  • Proposed Distillation with selective Input Gradient Regularization (DIGR) approach.
  • Utilized policy distillation and input gradient regularization for enhanced interpretability.
  • Evaluated DIGR on MiniGrid, Atari (Breakout), and CARLA Autonomous Driving.

Main Results:

  • DIGR produces interpretable saliency maps efficiently.
  • The approach improves the robustness of RL policies against adversarial attacks.
  • Experimental results demonstrate the effectiveness of DIGR across multiple tasks.

Conclusions:

  • DIGR offers a computationally efficient and interpretable solution for deep RL.
  • The method enhances RL policy robustness, making it suitable for real-world deployment.
  • DIGR represents a significant advancement in the interpretability and reliability of RL systems.