Related Experiment Video
Updated: Dec 17, 2025

Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
Fast generalization error bound of deep learning without scale invariance of activation functions
Yoshikazu Terada1, Ryoma Hirose2
1Graduate School of Engineering Science, Osaka University, 1-3 Machikaneyama-cho, Toyonaka, Osaka 560-8531, Japan; RIKEN Center for Advanced Intelligence Project (AIP), 1-4-1 Nihonbashi, Chuo-ku, Tokyo 103-0027, Japan.
Scale invariance of activation functions is not essential for fast deep learning convergence. A new analysis shows general activation functions achieve tight generalization error bounds, expanding Suzuki's framework.
Area of Science:
- Deep Learning Theory
- Machine Learning Analysis
- Neural Network Architectures
Background:
- Theoretical analysis of deep learning seeks to identify performance-driving features.
- Suzuki's (2018) generalization error framework suggests scale invariance in activation functions is crucial for tight error bounds.
- Common activation functions like sigmoid, tanh, and ELU lack scale invariance, potentially leading to slower convergence rates (O(1/n)).
Purpose of the Study:
- To derive a tight generalization error bound for deep neural networks with general activation functions, irrespective of scale invariance.
- To investigate whether scale invariance is truly essential for achieving fast convergence rates within Suzuki's (2018) theoretical framework.
- To demonstrate the broader applicability of Suzuki's (2018) framework to diverse activation functions in deep learning.
Main Methods:
- Applied Suzuki's (2018) generalization error analysis framework.
- Derived tight generalization error bounds for deep neural networks utilizing non-scale invariant activation functions.
- Theoretically analyzed convergence rates based on activation function properties.
Main Results:
- A tight generalization error bound was derived for deep neural networks with non-scale invariant activation functions, matching the quality of bounds derived under scale invariance.
- Demonstrated that scale invariance is not a prerequisite for achieving a fast rate of convergence within the analyzed framework.
- The derived bounds are essentially the same as those obtained by Suzuki (2018).
Conclusions:
- Scale invariance of activation functions is not essential for obtaining fast convergence rates in deep learning, contrary to previous interpretations of Suzuki's (2018) work.
- Suzuki's (2018) theoretical framework is applicable to a wider range of activation functions than previously assumed.
- This research broadens the understanding of deep learning generalization and the role of activation functions in theoretical analysis.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Improving Translational Accuracy
Improving Translational Accuracy
Propagation of Action Potentials
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
Graded Potential
Graded potentials fall into two categories: depolarizing and hyperpolarizing. Depolarizing graded potentials typically occur when sodium (Na+) or...