Related Experiment Video
Updated: Jan 17, 2026

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
Two-phase perspective on deep learning dynamics
Robert de Mello Koch1,2, Animik Ghosh1
1Huzhou University, School of Science, Huzhou 313000, China.
Deep neural network learning involves rapid curve-fitting then slower compression. This two-phase process, observed in grokking and double descent, is crucial for generalization.
Area of Science:
- Machine Learning
- Deep Learning Theory
- Computational Neuroscience
Background:
- Deep neural networks (DNNs) exhibit complex learning dynamics.
- Phenomena like grokking, double descent, and information bottleneck suggest non-trivial post-zero-training-error behavior.
- Understanding the temporal structure of DNN learning is key to improving generalization.
Purpose of the Study:
- To propose and validate a two-phase model of deep neural network learning: curve-fitting and compression.
- To demonstrate the shared temporal structure of grokking, double descent, and information bottleneck.
- To identify effective progress measures for DNN learning and explore the role of the compression phase.
Main Methods:
- Empirical analysis of DNN training dynamics in diverse settings.
- Measurement of mutual information between hidden layers and input data.
- Comparison of information-theoretic measures with circuit-based metrics (local complexity, linear mapping number).
Main Results:
- DNN learning consistently follows a two-phase trajectory: rapid initial fitting followed by slower compression.
- The timescales of grokking, double descent, and information bottleneck phenomena align empirically.
- Mutual information serves as a robust measure of learning progress, correlating with generalization.
Conclusions:
- Deep learning progresses through distinct curve-fitting and compression phases.
- The compression phase, analogous to renormalization group coarse-graining, is critical for generalization.
- Standard training may not optimize this compression phase, suggesting potential for algorithmic improvements.
Related Concept Videos
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
Observational Learning
Phase Transitions
Multi-input and Multi-variable systems
In the absence of...
Phase-lead and Phase-lag Controllers
