Related Experiment Video
Updated: Jul 26, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Information bottleneck theory of high-dimensional regression: relevancy, efficiency and optimality
Vudtiwat Ngampruetikorn1, David J Schwab1
1Initiative for the Theoretical Sciences, The Graduate Center, CUNY.
Researchers quantified overfitting in machine learning using residual information. Optimal algorithms minimize this noise, balancing relevant information for better predictions and understanding information efficiency.
Area of Science:
- Machine Learning
- Information Theory
- Statistical Modeling
Background:
- Overfitting is a significant challenge in machine learning, where models learn training data noise.
- Large neural networks often achieve zero training loss, creating a puzzling contradiction with overfitting.
- New approaches are needed to understand and quantify overfitting in complex models.
Purpose of the Study:
- To quantify overfitting using the concept of residual information.
- To investigate the trade-off between residual information (noise) and relevant information (predictive signals).
- To analyze the information efficiency of learning algorithms, particularly randomized ridge regression, compared to optimal algorithms.
Main Methods:
- Defined and quantified overfitting via residual information (bits encoding training data noise).
- Formulated an optimization problem to identify information-efficient learning algorithms.
- Solved the optimization for linear regression and compared results with randomized ridge regression.
- Applied random matrix theory to analyze high-dimensional linear map learning.
Main Results:
- Demonstrated a fundamental trade-off between residual and relevant information in machine learning models.
- Characterized the information efficiency of randomized ridge regression relative to theoretically optimal algorithms.
- Revealed the information complexity of learning linear maps in high dimensions.
- Identified information-theoretic analogs of double and multiple descent phenomena.
Conclusions:
- Residual information provides a principled way to quantify overfitting.
- Information efficiency offers a framework for designing better learning algorithms.
- Understanding information complexity is crucial for high-dimensional learning and explains phenomena like double descent.
Related Concept Videos
Column Efficiency: Rate Theory
During elution, a solute molecule experiences numerous transitions between stationary and mobile phases, exhibiting irregular residence times in...
Regression Toward the Mean
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
The Availability Heuristic
Outliers and Influential Points
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

