Related Experiment Video
Updated: Jun 21, 2025

08:32
Tracking Rats in Operant Conditioning Chambers Using a Versatile Homemade Video Camera and DeepLabCut
Published on: June 15, 2020
12.4K
Magnitude and angle dynamics in training single ReLU neurons
Sangmin Lee1, Byeongsu Sim1, Jong Chul Ye2
1Department of Mathematical Sciences, KAIST, Daejeon, Republic of Korea.
Summary
This study analyzes deep ReLU network training dynamics by examining single neuron weight vectors. We provide bounds on weight magnitude and angle to explain convergence, extending findings to multi-neuron networks.
Area of Science:
- Deep Learning
- Computational Neuroscience
- Optimization Theory
Background:
- Deep Rectified Linear Unit (ReLU) networks are crucial in deep learning.
- Understanding their training dynamics, particularly weight vector behavior, is essential but incomplete.
- Existing research lacks detailed analysis of single ReLU neuron weight dynamics.
Purpose of the Study:
- To investigate the training dynamics of gradient flow for single ReLU neurons.
- To analyze the convergence behavior by decomposing weight vectors into magnitude and angle.
- To extend these findings to multi-neuron networks and gradient descent.
Main Methods:
- Decomposition of weight vector gradient flow w(t) into magnitude ‖w(t)‖ and angle φ(t).
- Derivation of upper and lower bounds for these components.
- Empirical validation on two-layer multi-neuron networks.
- Generalization to gradient descent and experimental verification.
Main Results:
- Established theoretical bounds on the magnitude and angle of single ReLU neuron weight vectors.
- Demonstrated convergence dynamics through component analysis.
- Empirically extended findings to multi-neuron network settings.
- Confirmed theoretical results with gradient descent experiments.
Conclusions:
- The study provides a deeper understanding of deep ReLU network training dynamics at the single neuron level.
- The decomposition into magnitude and angle offers insights into convergence.
- Findings are applicable to broader network architectures and optimization algorithms.

