Neurocontrol for fixed-length trajectories in environments with soft barriers.

Michael Fairbank1, Danil Prokhorov2, David Barragan-Alcantar1

  • 1School of Computer Science and Electronic Engineering, University of Essex, Colchester, UK.

Summary

Training simulated agents with analytic policy gradients can lead to learning issues. Using soft barriers and gradient clipping or truncation helps avoid local minima and exploding gradients for better neurocontrol.