Related Experiment Video
Updated: Jan 5, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
On the Effect of the Activation Function on the Distribution of Hidden Nodes in a Deep Network
1Google Brain, Mountain View, CA 94043, U.S.A. plong@google.com.
Abstract:
We analyze the joint probability distribution on the lengths of the vectors of hidden variables in different layers of a fully connected deep network, when the weights and biases are chosen randomly according to gaussian distributions. We show that if the activation function satisfies a minimal set of assumptions, satisfied by all activation functions that we know that are used in practice, then, as the width of the network gets large, the "length process" converges in probability to a length map that is determined as a simple function of the variances of the random weights and biases and the activation function . We also show that this convergence may fail for that violate our assumptions. We show how to use this analysis to choose the variance of weight initialization, depending on the activation function, so that hidden variables maintain a consistent scale throughout the network.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Network Function of a Circuit
Transformations of Functions III
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Deactivation Processes: Jablonski Diagram