Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K
Decision Making: P-value Method01:09

Decision Making: P-value Method

6.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.8K
State Space Representation01:27

State Space Representation

499
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
499
Observational Learning01:12

Observational Learning

791
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
791
Time-Domain Interpretation of PD Control01:07

Time-Domain Interpretation of PD Control

348
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
348
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
1.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Analysis of gene expression patterns modulated by tuberculous pleural effusion-derived exosomal miRNAs in lung cancer.

Frontiers in genetics·2026
Same author

Thermally activated excess noise by subgap density-of-states in Si-doped ZnSnO thin-film transistor-type gas sensor.

Microsystems & nanoengineering·2026
Same author

A Multikinase Inhibitor AX-0085 Blocks FGFR1 Activation to Overcomes Osimertinib Resistance in Non-Small Cell Lung Cancer.

Biomedicines·2026
Same author

EGFR inhibitor suppresses tumor growth by blocking lipid uptake and cholesterol synthesis in non-small cell lung cancer.

Biochimica et biophysica acta. Molecular basis of disease·2025
Same author

Overview of Cervical Spine Injuries Caused by Diving Into Shallow Water on Jeju Island: A 9-Year Retrospective Study in a Regional Trauma Center.

Korean journal of neurotrauma·2025
Same author

Large-Scale Analysis of Defects in Atomically Thin Semiconductors using Hyperspectral Line Imaging.

Small (Weinheim an der Bergstrasse, Germany)·2024

Related Experiment Video

Updated: Jan 8, 2026

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
08:18

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control

Published on: August 15, 2020

5.4K

The Most Overestimated Q Value Regularization in High-Dimensional Discrete Action Spaces for Offline Reinforcement

Seunghwan Yu, Homin Park, Byungjin Ko

    IEEE Transactions on Neural Networks and Learning Systems
    |December 19, 2025
    PubMed
    Summary

    Deep reinforcement learning (DRL) for robotics faces challenges with data collection and Q-value overestimation. Our novel method, MQR, penalizes overestimated Q-values, improving stability and performance in high-dimensional action spaces.

    More Related Videos

    Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task
    11:18

    Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task

    Published on: June 1, 2015

    11.1K
    Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
    07:43

    Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios

    Published on: August 4, 2023

    2.6K

    Related Experiment Videos

    Last Updated: Jan 8, 2026

    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
    08:18

    WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control

    Published on: August 15, 2020

    5.4K
    Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task
    11:18

    Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task

    Published on: June 1, 2015

    11.1K
    Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
    07:43

    Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios

    Published on: August 4, 2023

    2.6K

    Area of Science:

    • Robotics
    • Artificial Intelligence
    • Machine Learning

    Background:

    • Deep reinforcement learning (DRL) is vital for robotic manipulation but hindered by data collection costs and risks.
    • Offline reinforcement learning (RL) trains on existing data but suffers from Q-value overestimation in high-dimensional discrete action spaces, impacting stability.
    • Out-of-distribution (OOD) actions increase rapidly in these spaces, exacerbating Q-value overestimation.

    Purpose of the Study:

    • To introduce a novel offline RL algorithm, Most Overestimated Q value Regularization (MQR), designed to mitigate Q-value overestimation.
    • To enhance training stability and prevent policy convergence errors in high-dimensional discrete action spaces for robotic manipulation.
    • To validate MQR's effectiveness in challenging robotic tasks with diverse environmental conditions.

    Main Methods:

    • Developed MQR, an offline RL algorithm that specifically penalizes the action with the most overestimated Q-value.
    • Implemented a targeted regularization strategy, differing from uniform penalties in existing methods.
    • Evaluated MQR on a robotic pushing and grasping task in simulated and real-world settings with varied object arrangements.

    Main Results:

    • MQR significantly outperformed baseline algorithms in robotic manipulation tasks.
    • Achieved a 96.94% clearance rate in simulations and 99.04% in real-world dense configurations.
    • Demonstrated high action efficiency and training stability, highlighting MQR's robustness and scalability.

    Conclusions:

    • MQR effectively mitigates Q-value overestimation in high-dimensional discrete action spaces for offline RL.
    • The algorithm shows strong performance, robustness, and adaptability for real-world robotic manipulation.
    • MQR holds significant potential for deployment in industrial robotics applications.