韦伯-费克纳定律在来自控制的时间差学习中作为推理.
Keiichiro Takahashi1, Taisuke Kobayashi2, Tomoya Yamanokuchi1
1Division of Information Science, Nara Institute of Science and Technology, Ikoma, Japan.
Frontiers in robotics and AI
|October 13, 2025
概括
本研究引入了一种非线性更新规则,用于强化学习 (RL),其灵感来源于生物学习. 韦伯-费克纳法 (WFL) 通过加快奖励获取和尽量减少惩罚来增强RL.
科学领域:
- 计算神经科学是一种神经科学.
- 机器学习 机器学习
- 人工智能的人工智能
背景情况:
- 标准的强化学习 (RL) 使用线性时间差 (TD) 错误更新,对待所有奖励均等.
- 生物系统在TD错误中表现出非线性,导致乐观或悲观的学习偏见.
- 这些非线性偏差是生物学习的潜在适应特征.
研究的目的:
- 探索一个理论框架,利用RL更新度和TD错误之间的非线性.
- 调查韦伯-费克纳定律 (WFL) 在控制作为推论框架内对RL的适用性.
- 通过奖励-惩罚系统来证明WFL在RL中的实际实用性.
主要方法:
- 对控制作为推论框架的分析,以确定RL中的非线性关系.
- 推导和应用韦伯-费克纳定律 (WFL) 来建模TD误差和更新大小之间的关系.
- 实施奖励-惩罚框架,以数量证明WFL对RL政策的影响.
主要成果:
- 确定了韦伯-费克纳定律 (WFL),描述了TD误差的感知如何随着值函数强度的变化而变化.
- 实行WFL证明了从低回报的情况加速逃脱.
- 实行WFL显示了对最低限度惩罚的加强追求.
结论:
- 拟议的RL算法结合WFL加速了奖励最大化,并有效地抑制了惩罚.
- 受生物学习启发的非线性更新规则在RL中提供了显著的优势.
- WFL提供了一个可行的机制,可以在人工学习系统中引入有益的偏见.
相关概念视频
Time-Domain Interpretation of PD Control
371
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
371
Feedback control systems
685
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
685
Frequency-Domain Interpretation of PD Control
351
Proportional-Derivative (PD) controllers are widely used in fan control systems to improve stability and performance. A fan control system can be effectively represented using a Bode plot to illustrate the impact of a PD controller through its transfer function. The Bode plot visually conveys how PD control modifies the fan's response across various frequencies, providing a frequency domain interpretation of the controller's behavior.
The proportional control gain, combined with the...
The proportional control gain, combined with the...
351
Time and frequency -Domain Interpretation of Phase-lag Control
389
Phase-lag controllers are widely used in control systems to improve stability and reduce steady-state errors. A dimmer switch controlling the brightness of a light bulb serves as a practical example of phase-lag control, gradually adjusting the bulb's brightness. Mathematically, phase-lag control or low-pass filtering is represented when the factor 'a' is less than 1.
Phase-lag controllers do not place a pole at zero, but instead influence the steady-state error by amplifying any...
Phase-lag controllers do not place a pole at zero, but instead influence the steady-state error by amplifying any...
389
Cognitive Learning
1.0K
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.0K
Transfer Function in Control Systems
1.5K
The transfer function is a fundamental concept in the analysis and design of linear time-invariant (LTI) systems. It offers a concise way to understand how a system responds to different inputs in the frequency domain. It serves as a bridge between the time-domain differential equations that describe system dynamics and the frequency-domain representation that facilitates easier manipulation and analysis.
To derive the transfer function, consider a general nth-order linear time-invariant...
To derive the transfer function, consider a general nth-order linear time-invariant...
1.5K


