关于超参数选择以保证RMSProp的收.
Jinlan Liu1, Dongpo Xu1, Huisheng Zhang2
1School of Mathematics and Statistics, Northeast Normal University, Changchun, 130024 China.
Cognitive neurodynamics
|December 23, 2024
概括
一个新的时间变化的RMSProp算法解决了深度学习中的融合问题. 这种优化方法,带有动态超参数,确保了各种目标的融合,并提高了基准数据集的性能.
科学领域:
- 深度学习优化优化
- 随机梯度下降变种 随机梯度下降变种
背景情况:
- 根平均平方传播 (RMSProp) 是深度学习中广泛使用的优化器.
- 最近的研究表明,即使在凸起的场景中,RMSProp也可能无法接近最佳解决方案.
研究的目的:
- 提出一个修改后的RMSProp算法,具有时间变化的超参数,以解决非融合问题.
- 理论分析拟议方法对凸和非凸问题的收性质.
主要方法:
- 为超参数引入一个时间变化的序列,而不是一个固定的值.
- 提供一个严格的数学证明,用于在非凸设置中的临界点的收,并具有特定的收率.
主要成果:
- 拟议的时间变化的RMSProp显示了对平滑,非凸起的目标的关键点的趋同.
- 建立了顺序的理论收率.
- 数字实验证实了随时间变化的RMSProp在基准数据集上的标准RMSProp的优势.
结论:
- 时间变化的RMSProp有效地解决了标准算法的分歧问题.
- 修改后的优化器为深度学习应用提供了更好的性能和理论保证.
- 这项工作为RMSProp的融合行为提供了更深入的理解.
相关概念视频
Time-Domain Interpretation of PD Control
83
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
83
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
40
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
40
Routh-Hurwitz Criterion II
182
In the application of the Routh-Hurwitz criterion, two specific scenarios can arise that complicate stability analysis.
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
The first scenario occurs when a singular zero appears in the first column of the Routh table. This situation creates a division by zero issues. To resolve this, a small positive or negative number, denoted as epsilon (∈), is substituted for the zero. The stability analysis proceeds by assuming a sign for ∈. If ∈ is positive, any sign change in the first...
182
Region of Convergence
377
The z-transform is a powerful mathematical tool used in the analysis of discrete-time signals and systems. It is a crucial tool in the analysis of discrete-time systems, but its convergence is limited to specific values of the complex variable z. This range of values, known as the Region of Convergence (ROC), is fundamental in determining the behavior and stability of a system or signal. The ROC defines the region in the complex plane where the z-transform converges, which can take various...
377
Region of Convergence of Laplace Tarnsform
484
The Region of Convergence (ROC) is a fundamental concept in signal processing and system analysis, particularly associated with the Laplace transform. The ROC represents an area in the complex plane where the Laplace transform of a given signal converges, determining the transform's applicability and utility.
Consider a decaying exponential signal that begins at a specific time. When deriving its Laplace transform, the time-domain variable is replaced with a complex variable. This...
Consider a decaying exponential signal that begins at a specific time. When deriving its Laplace transform, the time-domain variable is replaced with a complex variable. This...
484
Propagation of Uncertainty from Random Error
645
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
645


