在强化学习中实验量子加速
V Saggio1, B E Asenbeck2, A Hamann3
1University of Vienna, Faculty of Physics, Vienna Center for Quantum Science and Technology (VCQ), Vienna, Austria. valeria.saggio@univie.ac.at.
Nature
|March 11, 2021
概括
这项研究通过使用量子通信道加快代理学习来证明了强化学习的量子优势. 这一突破可以提高未来量子网络的人工智能效率.
科学领域:
- 人工智能
- 量子计算
- 量子通信
背景情况:
- 强化学习代理人通过环境互动和反来学习.
- 快速学习算法对于人工智能的发展至关重要.
- 以往使用量子力学来加速决策的尝试并没有缩短学习时间.
研究的目的:
- 通过量子力学来证明强化学习的加速.
- 通过结合量子和经典通信来评估改进.
- 在实用的纳米光子系统中实现和展示量子优势.
主要方法:
- 开发了一种使用量子通信通道的强化学习实验.
- 整合了一个快速主动反机制与纳米光子处理器.
- 用于量子通道接口的电信波长光子.
主要成果:
- 通过量子通信实现了智能体的学习速度提升.
- 通过结合量子和经典通信, 展示了学习进度的最佳控制.
- 在一个紧的,可调整的集成纳米光子处理器上实现协议.
结论:
- 量子通信道可以显著加速强化学习.
- 开发的纳米光子处理器为量子增强的人工智能提供了一个可扩展的平台.
- 这项工作为将量子优势整合到未来的通信网络铺平了道路.
相关概念视频
Reinforcement Schedules
316
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
316
Reinforcement
588
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
588
Observational Learning
606
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
606
Ampere-Maxwell's Law: Problem-Solving
917
A parallel-plate capacitor with capacitance C, whose plates have area A and separation distance d, is connected to a resistor R and a battery of voltage V. The current starts to flow at t = 0. What is the displacement current between the capacitor plates at time t? From the properties of the capacitor, what is the corresponding real current?
To solve the problem, we can use the equations from the analysis of an RC circuit and Maxwell's version of Ampère's law.
For the first part of the...
To solve the problem, we can use the equations from the analysis of an RC circuit and Maxwell's version of Ampère's law.
For the first part of the...
917
Randomized Experiments
8.6K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
8.6K
Reaction Quotient
51.6K
The status of a reversible reaction is conveniently assessed by evaluating its reaction quotient (Q). For a reversible reaction described by m A + n B ⇌ x C + y D, the reaction quotient is derived directly from the stoichiometry of the balanced equation as
51.6K


