Related Experiment Video
Updated: May 24, 2025

Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
Nearly Optimal Learning Using Sparse Deep ReLU Networks in Regularized Empirical Risk Minimization With Lipschitz
Ke Huang1, Mingming Liu2, Shujie Ma3
1Department of Statistics, University of California, Riverside, Riverside 92521, CA, U.S.A. khuan049@ucr.edu.
We introduce a Sparse Deep ReLU Network (SDRN) for regression problems. This novel estimator achieves near-optimal convergence rates, outperforming traditional networks by mitigating overfitting with fewer parameters.
Area of Science:
- Machine Learning
- Statistical Learning Theory
- Deep Learning
Background:
- Empirical risk minimization is a standard approach for statistical estimation.
- Lipschitz loss functions are commonly used in regression and classification.
- Deep neural networks often struggle with overfitting and parameter efficiency.
Purpose of the Study:
- To propose a Sparse Deep ReLU Network (SDRN) estimator for regression functions.
- To establish nonasymptotic excess risk bounds for the SDRN estimator.
- To analyze the convergence rate and parameter complexity of the SDRN.
Main Methods:
- Developing a sparse deep ReLU network architecture.
- Utilizing regularized empirical risk minimization with a Lipschitz loss function.
- Deriving nonasymptotic excess risk bounds for Sobolev spaces with mixed derivatives.
Main Results:
- The SDRN estimator achieves a nearly optimal minimax convergence rate, comparable to 1D nonparametric regression.
- The convergence rate is logarithmic in feature dimension when fixed, and slightly slower when dimension grows with sample size.
- The depth of the SDRN grows logarithmically with sample size, while nodes/weights grow polynomially.
Conclusions:
- The proposed SDRN estimator effectively estimates regression functions.
- SDRN overcomes overfitting issues in conventional feedforward networks.
- SDRN offers improved depth and parameter efficiency for enhanced regression performance.
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Residual Plots
When the residual values are plotted against the variable x, it is called a residual...
Empirical Method to Interpret Standard Deviation
This rule is used widely in statistics to calculate the proportion of data values...
Purposive Learning

