Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Woodward–Hoffmann Selection Rules and Microscopic Reversibility01:34

Woodward–Hoffmann Selection Rules and Microscopic Reversibility

3.2K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.2K
Decision Making: P-value Method01:09

Decision Making: P-value Method

5.5K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.5K
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

133
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
133
Observational Learning01:12

Observational Learning

225
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
225
Reinforcement01:23

Reinforcement

290
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
290
Decision Making01:20

Decision Making

152
Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
152

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Spatiotemporal distribution of legacy and alternative per- and Polyfluoroalkyl substances (PFASs) in major rivers of the Pearl river delta.

Journal of environmental sciences (China)·2026
Same author

Associations of legacy and emerging per- and polyfluoroalkyl substances (PFAS) with aquatic communities in a typical subtropical estuary.

Environment international·2026
Same author

Spatiotemporal variation of marine microbes in the Taiwan strait ecosystem.

Environmental research·2025
Same author

Observations and potential source regions of HFC-152a in southeastern China.

Environmental research·2025
Same author

Multiple impacts of human activities on environmental fate of per- and polyfluoroalkyl substances (PFAS) in the Xiaoqing River of China.

Environmental pollution (Barking, Essex : 1987)·2025
Same author

Metabolomic analysis reveals contrasting effects of PFOS and PFAS on cyanobacterial bloom and metabolic pathways in eutrophic water.

Environmental pollution (Barking, Essex : 1987)·2025

Related Experiment Video

Updated: Jul 26, 2025

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
11:54

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

Published on: May 8, 2021

4.5K

Differentiable Logic Policy for Interpretable Deep Reinforcement Learning: A Study From an Optimization Perspective.

Xin Li, Haojie Lei, Li Zhang

    IEEE Transactions on Pattern Analysis and Machine Intelligence
    |June 13, 2023
    PubMed
    Summary

    This study introduces Differentiable Inductive Logic Programming (DILP) for interpretable Deep Reinforcement Learning (DRL). Mirror Descent for Policy Optimization (MDPO) effectively addresses constraints in DILP-based policies, enhancing interpretability.

    More Related Videos

    Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
    07:05

    Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

    Published on: September 10, 2018

    6.0K
    A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
    05:41

    A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

    Published on: February 6, 2020

    9.5K

    Related Experiment Videos

    Last Updated: Jul 26, 2025

    Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
    11:54

    Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface

    Published on: May 8, 2021

    4.5K
    Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
    07:05

    Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents

    Published on: September 10, 2018

    6.0K
    A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
    05:41

    A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

    Published on: February 6, 2020

    9.5K

    Area of Science:

    • Artificial Intelligence
    • Machine Learning
    • Deep Reinforcement Learning

    Background:

    • Policy interpretability is a significant challenge in Deep Reinforcement Learning (DRL).
    • Current DRL methods often lack transparency, hindering understanding and trust.
    • Representing policies using symbolic methods could improve interpretability.

    Purpose of the Study:

    • To explore interpretable Deep Reinforcement Learning (DRL) by representing policies with Differentiable Inductive Logic Programming (DILP).
    • To provide a theoretical and empirical analysis of DILP-based policy learning from an optimization viewpoint.
    • To introduce a novel optimization approach for DILP-based policies.

    Main Methods:

    • Representing DRL policies using Differentiable Inductive Logic Programming (DILP).
    • Formulating DILP-based policy learning as a constrained policy optimization problem.
    • Proposing and analyzing Mirror Descent for Policy Optimization (MDPO) to handle DILP policy constraints.
    • Deriving closed-form regret bounds for MDPO with function approximation.
    • Investigating the convexity of DILP-based policies.

    Main Results:

    • Identified DILP-based policy learning as a constrained optimization problem.
    • Developed Mirror Descent for Policy Optimization (MDPO) as an effective solution for DILP policy constraints.
    • Derived theoretical regret bounds for MDPO, aiding DRL framework design.
    • Empirical results validated the effectiveness of MDPO and its on-policy variant compared to mainstream methods.
    • Demonstrated the benefits of MDPO through the study of DILP-based policy convexity.

    Conclusions:

    • Differentiable Inductive Logic Programming (DILP) offers a promising approach for interpretable Deep Reinforcement Learning (DRL).
    • Mirror Descent for Policy Optimization (MDPO) provides a theoretically sound and empirically effective method for optimizing constrained DILP-based policies.
    • The proposed framework enhances policy interpretability while maintaining competitive performance in DRL tasks.