具有模拟memristor的演员关键网络模拟基于奖励的学习
Kevin Portner1, Till Zellweger1, Flavio Martinelli2
1Integrated Systems Laboratory, ETH Zurich, Zurich, Switzerland.
Nature machine intelligence
|December 24, 2025
概括
研究人员开发了一种新的生物灵感计算系统,使用模拟memristors进行强化学习. 这种硬件完全在内存中进行在线训练和动作选择,推进了神经形态计算引擎.
科学领域:
- 神经形态工程的神经形态工程
- 计算神经科学是一种神经科学.
背景情况:
- 当前的生物启发计算通常仅模仿大脑部分.
- 记忆硬件通常用于学习算法的有限功能,需要软件来完成复杂的任务.
研究的目的:
- 使用模拟记忆器演示一个完全集成的强化学习系统.
- 创建一个生物灵感的神经网络架构,在硬件上完全执行基于奖励的学习.
主要方法:
- 在模拟memristors上实现了一个演员关键时间差算法.
- 使用memristors作为多用途元件:突触权重,权重更新计算器和动作决定器.
- 在T-迷宫和莫里斯水迷宫导航任务中测试了框架.
主要成果:
- 实现了在线体重训练和直接硬件计算时间差错.
- 启用了完整的内存计算,消除了用于体重训练的数据移动.
- 通过使用基于memristor的系统在模拟环境中成功演示导航.
结论:
- 这项工作介绍了第一个完全内存的在线强化学习系统,用于生物启发的计算.
- 这种方法为更高效和类似大脑的神经形态计算引擎铺平了道路.
- 模拟记忆器可以作为先进的人工智能硬件的多功能组件.
相关概念视频
Observational Learning
782
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
782
Design Example: Frog Muscle Response
544
A student is tasked to work on an intriguing experiment involving an RL (Resistor-Inductor) circuit to study the muscle response of a frog's leg to electrical stimulation. The RL circuit plays a crucial role in this experiment, providing the means to control and measure the electrical impulses that trigger muscle contraction.
When the switch connecting the RL circuit is closed, a brief muscle contraction is observed. This is because, at a steady state, the inductor acts like a short...
When the switch connecting the RL circuit is closed, a brief muscle contraction is observed. This is because, at a steady state, the inductor acts like a short...
544


