Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model01:13

Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model

182
Drugs administered through various routes can lead to nonlinear elimination, resulting in complex pharmacokinetic behaviors crucial to understanding efficacious drug dosing.
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
182
Alternative Sets of Equilibrium Equations01:31

Alternative Sets of Equilibrium Equations

791
When analyzing the behavior of structures, engineers often rely on the concept of equilibrium. This refers to the state where all forces and moments acting on a system balance each other, resulting in no net movement or rotation. In many cases, equilibrium can be described by a set of standard equations. However, in some situations, alternative sets of equilibrium equations must be used to describe the system's behavior accurately.
One example of such a situation can be observed in a...
791
Observational Learning01:12

Observational Learning

655
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
655
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

273
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
273
Maxam-Gilbert Sequencing01:05

Maxam-Gilbert Sequencing

12.0K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
12.0K
Woodward–Hoffmann Selection Rules and Microscopic Reversibility01:34

Woodward–Hoffmann Selection Rules and Microscopic Reversibility

3.5K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Hybrid Event-Triggered Tracking Control With Critic Learning for Nonlinear Networked Systems.

IEEE transactions on cybernetics·2026
Same author

Tacit mechanism: Bridging pre-training of individuality to multi-agent adversarial coordination.

Neural networks : the official journal of the International Neural Network Society·2025
Same author

Balancing State Exploration and Skill Diversity in Unsupervised Skill Discovery.

IEEE transactions on cybernetics·2025
Same author

Last-Iterate Convergence to Approximate Nash Equilibria in Multiplayer Imperfect Information Games.

IEEE transactions on neural networks and learning systems·2025
Same author

Meta Learning Task Representation in Multiagent Reinforcement Learning: From Global Inference to Local Inference.

IEEE transactions on neural networks and learning systems·2025
Same author

Plinabulin exerts an anti-proliferative effect via the PI3K/AKT/mTOR signaling pathways in glioblastoma.

Iranian journal of basic medical sciences·2025

Related Experiment Video

Updated: Nov 26, 2025

The HoneyComb Paradigm for Research on Collective Human Behavior
06:48

The HoneyComb Paradigm for Research on Collective Human Behavior

Published on: January 19, 2019

9.7K

Online Minimax Q Network Learning for Two-Player Zero-Sum Markov Games.

Yuanheng Zhu, Dongbin Zhao

    IEEE Transactions on Neural Networks and Learning Systems
    |December 11, 2020
    PubMed
    Summary

    Researchers developed a new method using deep reinforcement learning (DRL) to find Nash equilibrium policies in complex games. This approach combines game theory and dynamic programming for efficient online learning in two-player zero-sum Markov games.

    More Related Videos

    Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
    05:30

    Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit

    Published on: September 8, 2023

    948
    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
    08:18

    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control

    Published on: August 15, 2020

    5.2K

    Related Experiment Videos

    Last Updated: Nov 26, 2025

    The HoneyComb Paradigm for Research on Collective Human Behavior
    06:48

    The HoneyComb Paradigm for Research on Collective Human Behavior

    Published on: January 19, 2019

    9.7K
    Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
    05:30

    Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit

    Published on: September 8, 2023

    948
    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
    08:18

    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control

    Published on: August 15, 2020

    5.2K

    Area of Science:

    • Game Theory
    • Artificial Intelligence
    • Reinforcement Learning

    Background:

    • The Nash equilibrium is a key concept in game theory, representing a state of minimal exploitability between opponents.
    • Two-player zero-sum Markov games (TZMGs) present complex strategic decision-making challenges.

    Purpose of the Study:

    • To develop and validate an online learning algorithm for finding Nash equilibrium policies in TZMGs.
    • To integrate deep reinforcement learning (DRL) with generalized policy iteration (GPI) for solving TZMGs.

    Main Methods:

    • Formulated the TZMG problem as a Bellman minimax equation.
    • Applied generalized policy iteration (GPI) with neural networks for Q-function approximation.
    • Developed an online minimax Q-network learning algorithm incorporating experience replay, dueling networks, and double Q-learning.

    Main Results:

    • Successfully combined DRL techniques with GPI to determine TZMG Nash equilibrium.
    • Proved the convergence of the online learning algorithm, offering insights for single-agent problems.
    • Experimental validation demonstrated the algorithm's effectiveness on various TZMG examples.

    Conclusions:

    • The proposed DRL-based approach offers a novel and effective method for solving TZMGs.
    • The convergence proof extends theoretical understanding to both multi-agent and single-agent reinforcement learning.
    • This work advances the application of AI in strategic decision-making and game theory.