Related Experiment Video
Updated: Sep 20, 2025

Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
A New Accelerated Off-Policy Stochastic Preconditioned TD(0) Algorithm
None:
In this article, we consider policy evaluation in off-policy reinforcement learning and propose a novel procedure (Stochastic Preconditioned Temporal Difference (SPTD)) that achieves the optimal convergence rate under linear function approximation. The procedure has a linear computational complexity of the dimension of the feature space in each iteration. Under Markovian sampling, we establish finite-sample rates when the target policy can be different from the behavior policy for data generation. Our procedure is the first algorithm for the off-policy policy evaluation that has the optimal rate $\mathcal {O}(1/t)$O(1/t) under the mean square error. We also provide the first result on the asymptotic distribution and give the nearly optimal step size $\alpha _{t} = \mathcal {O}(t^{-2/3})$αt=O(t-2/3). The numerical performance of the procedure is studied in both on-policy and off-policy settings. Extensive numerical experiments demonstrate that our procedure uniformly outperforms existing methods.
Related Concept Videos
Fast Decoupled and DC Powerflow
Statically Indeterminate Problem Solving
Time-Domain Interpretation of PD Control
Consider the example of control of motor torque. Initially, a positive...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
PD Controller: Design
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Bernoulli's Equation: Problem Solving
The first step is to compute the cross-sectional areas of the pipe and the Venturi throat to analyze the pressure difference indicated by the pressure gauge. Next, the continuity...

