Related Experiment Video
Updated: Oct 22, 2025

Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
Non-differentiable saddle points and sub-optimal local minima exist for deep ReLU networks
Bo Liu1, Zhaoying Liu1, Ting Zhang1
1College of Computer Science, Faculty of Information Technology, Beijing University of Technology, Beijing, China.
Sub-optimal local minima and saddle points are proven to exist in deep ReLU networks. This finding impacts optimization algorithms for deep learning models, especially with cross-entropy loss.
Area of Science:
- Deep Learning
- Optimization Theory
- Computational Mathematics
Background:
- The performance of deep neural networks is significantly influenced by the characteristics of their loss landscape, particularly the presence of local minima and saddle points.
- Understanding these landscape features is crucial for developing effective optimization algorithms in deep learning.
Purpose of the Study:
- To theoretically investigate the existence of non-differentiable sub-optimal local minima and saddle points in deep Rectified Linear Unit (ReLU) networks of arbitrary depth.
- To analyze the implications of these landscape features on optimization algorithms used for training deep learning models.
Main Methods:
- Theoretical analysis of the loss surface for deep ReLU networks.
- Mathematical proofs for the existence of non-differentiable saddle points and sub-optimal local minima under specific loss functions (squared loss, cross-entropy loss).
- Empirical validation using experiments on both real-world and synthetic datasets.
Main Results:
- Existence of non-differentiable saddle points is proven for deep ReLU networks with squared loss or cross-entropy loss under standard assumptions.
- Non-differentiable sub-optimal local minima are proven to exist in deep ReLU networks with cross-entropy loss when certain data distribution conditions are met.
- Experimental results align with theoretical predictions, confirming the presence of these critical points in the loss landscape.
Conclusions:
- The loss landscape of deep ReLU networks is characterized by the guaranteed existence of non-differentiable saddle points and, under certain conditions, sub-optimal local minima.
- These findings provide theoretical insights into the challenges faced by optimization algorithms in deep learning and suggest avenues for algorithm improvement.
- The study underscores the importance of considering the complex geometry of the loss landscape for effective deep neural network training.
More Related Videos
03:31Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
06:45Design and Application of a Fault Detection Method Based on Adaptive Filters and Rotational Speed Estimation for an Electro-Hydrostatic Actuator
Published on: October 28, 2022
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
Survival Tree
Building a Survival Tree
Constructing a...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Plotting and Calibrating the Root Locus
The maximum gain occurs at the breakaway points between open-loop poles on the real axis, while the minimum gain is...