Related Experiment Video
Updated: Sep 23, 2025

Author Spotlight: Development of a Minimally Invasive Large-Animal Model for Reliable and Reproducible Cardiovascular Research
Published on: October 20, 2023
Anti-Martingale Proximal Policy Optimization
Abstract:
Since the sample data after one exploration process can only be used to update network parameters once in on-policy deep reinforcement learning (DRL), a high sample efficiency is necessary to accelerate the training process of on-policy DRL. In the proposed method, a submartingale criterion is proposed on the basis of the equivalence relationship between the optimal policy and martingale, and then an advanced value iteration (AVI) method is proposed to conduct value iteration with a high accuracy. Based on this foundation, an anti-martingale (AM) reinforcement learning framework is established to efficiently select the sample data that is conducive to policy optimization. In succession, an AM proximal policy optimization (AMPPO) method, which combines the AM framework with proximal policy optimization (PPO), is proposed to reasonably accelerate the updating process of state value that satisfies the submartingale criterion. Experimental results on the Mujoco platform show that AMPPO can achieve better performance than several state-of-the-art comparative DRL methods.
More Related Videos
Related Concept Videos
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Time-Domain Interpretation of PD Control
Consider the example of control of motor torque. Initially, a positive...
Field Procedure for Staking Out Curves
Mitral Stenosis III: Medical Management
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...

