Reinforcement
Reinforcement Schedules
Observational Learning
Associative Learning
Introduction to Learning
Avoidance Learning and Learned Helplessness
您也可能阅读
通过共同作者、期刊和引用图与本文相关的文章。
Soumia Mehimeh1, Xianglong Tang1
1School of Computer Science and Technology, Harbin Institute of Technology, 92 West Dazhi Street, Nangang District, Harbin 150001, China.
本研究介绍了LogQInit,这是一种用于强化学习中的价值函数转移的新方法. 它利用价值函数的日志常态分布,在任务变化的稀疏奖励环境中提高性能.
科学领域:
背景情况:
研究的目的:
主要方法:
主要成果:
结论: