Related Experiment Video
Updated: Sep 28, 2025

Using a Split-belt Treadmill to Evaluate Generalization of Human Locomotor Adaptation
Published on: August 23, 2017
Calibrating the Adaptive Learning Rate to Improve Convergence of ADAM
Qianqian Tong1, Guannan Liang1, Jinbo Bi1
1Computer Science and Engineering, University of Connecticut, Storrs, CT 06269.
Abstract:
Adaptive gradient methods (AGMs) have become popular in optimizing the nonconvex problems in deep learning area. We revisit AGMs and identify that the adaptive learning rate (A-LR) used by AGMs varies significantly across the dimensions of the problem over epochs (i.e., anisotropic scale), which may lead to issues in convergence and generalization. All existing modified AGMs actually represent efforts in revising the A-LR. Theoretically, we provide a new way to analyze the convergence of AGMs and prove that the convergence rate of Adam also depends on its hyper-parameter є, which has been overlooked previously. Based on these two facts, we propose a new AGM by calibrating the A-LR with an activation (softplus) function, resulting in the Sadam and SAMSGrad methods. We further prove that these algorithms enjoy better convergence speed under nonconvex, non-strongly convex, and Polyak-Łojasiewicz conditions compared with Adam. Empirical studies support our observation of the anisotropic A-LR and show that the proposed methods outperform existing AGMs and generalize even better than S-Momentum in multiple deep learning tasks.
Related Concept Videos
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...
Improving Translational Accuracy
Calibration Curves: Correlation Coefficient
Regression Toward the Mean
Instrument Calibration
Analytical Balance Calibration
An analytical balance measures mass and requires regular calibration to...
Load-frequency control

