通过D3QN增强DBR镜像设计:一种强化学习方法
Seungjun Yu1, Haneol Lee1, Changyoung Ju1
1Department of Electrical Engineering, Pohang University of Science and Technology (POSTECH), Pohang, Republic of Korea.
PloS one
|August 22, 2024
概括
使用D3QN进行深度强化学习,优化分布式布拉格反射器 (DBRs) 以提高反射率和紧性. 这种新的方法显著优于传统的光学设计技术.
科学领域:
- 光学工程是指光学工程.
- 人工智能的人工智能
- 材料科学 材料科学 材料科学
背景情况:
- 光学系统对于现代电子和通信至关重要.
- 像分布式布拉格反射器 (DBR) 这样的光学元件的传统设计方法是模拟密集型和耗时的.
- 光学系统设计的创新是由提高性能和小型化的需求驱动的.
研究的目的:
- 为设计分布式布拉格反射器 (DBRs) 引入一种新的深度强化学习方法.
- 优化DBR的多层结构,以提高反射率和减少尺寸.
- 将拟议方法的效率和性能与传统技术进行比较.
主要方法:
- 使用深度强化学习算法D3QN (决斗架构和双重Q网络).
- 应用D3QN来优化分布式布拉格反射器的多层结构.
- 用D3QN设计的DBR与使用转移矩阵方法 (TMM) 设计的DBR进行了比较.
主要成果:
- 与TMM衍生设计相比,使用D3QN设计的DBR显示出20.5%更高的反射率.
- 使用D3QN设计的DBR的大小减少了61.2%.
- 在传统的代模拟方法中,D3QN方法显示出更高的效率.
结论:
- 深度强化学习,特别是D3QN方法,为光学设计提供了一个有希望和高效的替代方案.
- D3QN方法显著提高了DBR的反射性能和紧性.
- 未来的工作可以将D3QN扩展到复杂的2D和3D光学设计结构.
相关概念视频
Reinforcement
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...


