Related Experiment Video
Updated: Nov 27, 2025

Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
Mutual Information Based Learning Rate Decay for Stochastic Gradient Descent Training of Deep Neural Networks
1IBM Research, Bangalore 560045, India.
Abstract:
This paper demonstrates a novel approach to training deep neural networks using a Mutual Information (MI)-driven, decaying Learning Rate (LR), Stochastic Gradient Descent (SGD) algorithm. MI between the output of the neural network and true outcomes is used to adaptively set the LR for the network, in every epoch of the training cycle. This idea is extended to layer-wise setting of LR, as MI naturally provides a layer-wise performance metric. A LR range test determining the operating LR range is also proposed. Experiments compared this approach with popular alternatives such as gradient-based adaptive LR algorithms like Adam, RMSprop, and LARS. Competitive to better accuracy outcomes obtained in competitive to better time, demonstrate the feasibility of the metric and approach.
Related Concept Videos
Current Growth And Decay In RL Circuits
Survival Tree
Building a Survival Tree
Constructing a...
Neural Regulation
Improving Translational Accuracy
Improving Translational Accuracy
Resting Potential Decay
At rest, the K+ is the main ion that moves across the membrane...