内核形状重规范化解释了有限贝叶斯单层隐藏网络中的输出-输出相关性
P Baglioni1, L Giambagli2, A Vezzani3,4,5
1INFN, sezione di Milano Bicocca, Piazza della scienza 3, 20126 Milano, Italy.
Physical review. E
|August 1, 2025
概括
有限宽度神经网络显示出输出相关性,而无限宽度模型中没有这种相关性. 这项研究用贝叶斯深度学习的内核形状重规范化来解释这些相关性,并通过数值实验验证.
科学领域:
- 机器学习 机器学习
- 深度学习理论 深度学习理论
- 贝叶斯的推理是贝叶斯的推理.
背景情况:
- 有限宽度的神经网络表现出复杂的输出-输出相关性,特别是多个读出神经元.
- 这些相关性在无限宽度极限中消失,这种现象被称为惰训练.
- 了解这些有限宽度效应对于深度学习的完整理论至关重要.
研究的目的:
- 为了合理化在有限宽度神经网络中对非微不足道的输出-输出相关性的经验观察.
- 利用贝叶斯深度学习的比例极限来解释这些相关性.
- 用数值实验验证理论框架.
主要方法:
- 使用贝叶斯深度学习的比例极限 (P/N有限),其中P是训练集大小,N是网络宽度.
- 为神经网络高斯过程 (NNGP) 内核开发一个内核形状重规范化理论.
- 进行数值实验以评估概括和量化输出-输出相关性.
主要成果:
- 有限网络中的输出-输出相关性是由重新规范化的NNGP内核解释的.
- 比例极限为理解这些有限宽度现象提供了一个理论框架.
- 数字实验在数量上与对应关系的理论预测相匹配.
结论:
- 核心形状的重新规范化是理解有限贝叶斯一层隐藏网络的相关性的关键.
- 比率极限中的贝叶斯深度学习框架准确地捕捉了有限宽度的效应.
- 这项工作弥合了无限宽度理论和有限宽度实证观测之间的差距.
相关概念视频
Reducing Line Loss
194
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
194
Residuals and Least-Squares Property
7.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.8K
Coefficient of Correlation
6.4K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable x and the dependent variable y.
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
If you suspect a linear relationship between x and y, then r can measure how strong the linear relationship is.
What the VALUE of r tells us:
The value of r is always between –1 and +1: –1 ≤ r ≤ 1.
The size of the correlation r indicates the...
6.4K
Normal Distribution
12.5K
The normal, a continuous distribution, is the most important of all the distributions. Its graph is a bell-shaped symmetrical curve, which is observed in almost all disciplines. Some of these include psychology, business, economics, the sciences, nursing, and, of course, mathematics. Some instructors may use the normal distribution to help determine students’ grades. Most IQ scores are normally distributed. Often real-estate prices fit a normal distribution. The normal distribution is...
12.5K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
717
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
717
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K

