Related Experiment Video
Updated: Feb 7, 2026

Fabrication of 1-D Photonic Crystal Cavity on a Nanofiber Using Femtosecond Laser-induced Ablation
Published on: February 25, 2017
Scalable photonic reinforcement learning by time-division multiplexing of laser chaos
Makoto Naruse1, Takatomo Mihana2, Hirokazu Hori3
1Network System Research Institute, National Institute of Information and Communications Technology, 4-2-1 Nukui-kita, Koganei, Tokyo, 184-8795, Japan. naruse@nict.go.jp.
Abstract:
Reinforcement learning involves decision-making in dynamic and uncertain environments and constitutes a crucial element of artificial intelligence. In our previous work, we experimentally demonstrated that the ultrafast chaotic oscillatory dynamics of lasers can be used to efficiently solve the two-armed bandit problem, which requires decision-making concerning a class of difficult trade-offs called the exploration-exploitation dilemma. However, only two selections were employed in that research; hence, the scalability of the laser-chaos-based reinforcement learning should be clarified. In this study, we demonstrated a scalable, pipelined principle of resolving the multi-armed bandit problem by introducing time-division multiplexing of chaotically oscillated ultrafast time series. The experimental demonstrations in which bandit problems with up to 64 arms were successfully solved are presented where laser chaos time series significantly outperforms quasiperiodic signals, computer-generated pseudorandom numbers, and coloured noise. Detailed analyses are also provided that include performance comparisons among laser chaos signals generated in different physical conditions, which coincide with the diffusivity inherent in the time series. This study paves the way for ultrafast reinforcement learning by taking advantage of the ultrahigh bandwidths of light wave and practical enabling technologies.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Long Division of Polynomials
Reinforcements in Concrete
Corrosion of Reinforcement
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
Reinforcement Schedules
Once a behavior is learned,...
Cranial Part of Parasympathetic Division
The vagus nerve (cranial nerve X) alone accounts for approximately 75...

