Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reinforcement Schedules01:24

Reinforcement Schedules

Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Role of Shaping in Operant Conditioning01:19

Role of Shaping in Operant Conditioning

Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
Law of Effect01:06

Law of Effect

B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Observational Learning01:12

Observational Learning

Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Purposive Learning01:22

Purposive Learning

E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a bonus...
Operant Conditioning Intervention01:24

Operant Conditioning Intervention

Operant conditioning serves as a foundational principle in therapeutic interventions aimed at modifying maladaptive behaviors. Central to this approach is the notion that behaviors, both adaptive and maladaptive, are learned through reinforcement. By analyzing the environmental factors that reinforce problematic behaviors, clinicians can design interventions to weaken these reinforcements and replace maladaptive behaviors with healthier alternatives.
In operant conditioning, behaviors that are...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The Felodipine Event Reduction (FEVER) Study: a randomized long-term placebo-controlled trial in Chinese hypertensive patients.

Journal of hypertension·2005
Same author

More tea for septic patients?--Green tea may reduce endotoxin-induced release of high mobility group box 1 and other pro-inflammatory cytokines.

Medical hypotheses·2005
Same author

[Cytokines secretion by peripheral blood mononuclear cells from hepatitis C patients after stimulation with synthetic peptides at the highly variable region].

Zhonghua shi yan he lin chuang bing du xue za zhi = Zhonghua shiyan he linchuang bingduxue zazhi = Chinese journal of experimental and clinical virology·2005
Same author

Archaeal proteasomes and other regulatory proteases.

Current opinion in microbiology·2005
Same author

Effect of carbamate esters on neurite outgrowth in differentiating human SK-N-SH neuroblastoma cells.

Chemico-biological interactions·2005
Same author

Resolving overlapped spectra with curve fitting.

Spectrochimica acta. Part A, Molecular and biomolecular spectroscopy·2005

Related Experiment Videos

Interpretable reinforcement learning with structured policies for prescriptive supply chain analytics.

Yixuan Huang1, Wei Li2

  • 1Faculty of Transportation Engineering, Kunming University of Science and Technology, Jingming South Road, Kunming, 650500, China.

Scientific Reports
|July 15, 2026
PubMed
Summary

Structured Policy Reinforcement Learning (SPRL) makes reinforcement learning (RL) interpretable for supply chain management. This approach reduces costs by 45% while improving service levels, enabling trustworthy autonomous systems.

Keywords:
Decision treesInterpretable reinforcement learningInventory controlPolicy distillationSupply chain analytics

Related Experiment Videos

Area of Science:

  • Operations Research
  • Artificial Intelligence
  • Supply Chain Management

Background:

  • Reinforcement Learning (RL) shows promise for complex supply chain decisions like inventory control.
  • Deep RL methods are often "black boxes," hindering trust and adoption in industrial settings.
  • Interpretability remains a significant challenge for applying RL in high-stakes environments.

Purpose of the Study:

  • To bridge the interpretability-performance gap in RL for supply chain management.
  • To introduce a novel framework, Structured Policy Reinforcement Learning (SPRL), for transparent RL.
  • To enable trustworthy autonomous decision-making in real-world supply chains.

Main Methods:

  • Developed SPRL, a framework embedding transparency into the RL agent's learning process.
  • Utilized a decoupled architecture to distill Q-learning agent value estimates into a Decision Tree (DT).
  • Incorporated a hard Inventory Position Cap constraint to stabilize training and enforce lean policies.

Main Results:

  • The SPRL-DT policy achieved a mean total cost reduction of 45% compared to the constrained DQN baseline.
  • Maintained an excellent service level exceeding 91%.
  • Resulted in a low-complexity, interpretable structure with [Formula: see text] nodes.

Conclusions:

  • SPRL offers a transparent and verifiable solution for RL in supply chain management.
  • Demonstrates the potential for trustworthy autonomous systems through interpretable RL.
  • Paves the way for practical deployment of RL in dynamic inventory control and other supply chain operations.