Related Experiment Video
Updated: Jan 11, 2026

Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
Effective Learning Rules as Natural Gradient Descent
Lucas Shoji1, Kenta Suzuki2, Leo Kozachkov3,4
1Department of Physics and Department of Brain and Cognitive Sciences, MIT, Cambridge, MA 02139, USA lshoji@mit.edu.
Abstract:
We establish that a broad class of effective learning rules-those that improve a scalar performance measure over a given time window-can be expressed as natural gradient descent with respect to an appropriately defined metric. Specifically, parameter updates in this class can always be written as the product of a symmetric positive-definite matrix and the negative gradient of a loss function encoding the task. Given the high level of generality, our findings formally support the idea that the gradient is a fundamental object underlying all learning processes. Our results are valid across a wide range of common settings, including continuous- time, discrete-time, stochastic, and higher-order learning rules, as well as loss functions with explicit time dependence. Beyond providing a unified framework for learning, our results also have practical implications for control as well as experimental neuroscience.
Related Concept Videos
Gradient and Del Operator
Regression Toward the Mean
Observational Learning
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Purposive Learning
What is Natural Selection?
