Related Experiment Videos
In situ trainable optical convolution enabled by an AWG core and gradient-based model-free optimization
Optics Express
|August 14, 2026
Summary
We developed an optical tensor processing unit (OTPU) using an arrayed waveguide grating (AWG) for efficient AI acceleration. This system achieves high accuracy in image classification and edge extraction tasks, offering a scalable pathway for photonic neural networks.
Area of Science:
- Photonics
- Artificial Intelligence
- Optical Computing
Background:
- Photonic neural networks (PNNs) offer potential for high-speed, low-power computation.
- Existing PNNs face challenges in scalability and efficient training, especially with physical non-idealities.
Purpose of the Study:
- To propose and validate a novel optical tensor processing unit (OTPU) architecture.
- To develop an effective training framework for PNNs leveraging optical components.
Main Methods:
- Utilized an arrayed waveguide grating (AWG) for parallel multiply-accumulate (MAC) operations.
- Implemented a gradient-based model-free optimization (G-MFO) framework for training.
- Encoded signed weights using Mach-Zehnder modulators (MZMs) and wavelength tuning.
Main Results:
- Achieved 98.80% accuracy on MNIST classification with 7.40 effective number of bits (ENOB) precision.
- Demonstrated versatile edge extraction on handwritten digits and breast ultrasound images.
- Showcased scalability through the use of additional wavelength channels.
Conclusions:
- The proposed AWG-based OTPU, trained with G-MFO, provides a scalable and training-efficient solution for PNN accelerators.
- This co-designed system addresses practical challenges in photonic computing for AI applications.
- The approach enables efficient in-situ training of non-differentiable optical weights.
Related Concept Videos
Gradient Vectors and Their Applications
Every point on a topographical map corresponds to a particular elevation, so the landscape can be modeled as a surface whose height depends on horizontal position. From any given location, a hiker may face infinitely many directions, but only one direction produces the fastest possible increase in elevation. This unique route is called the direction of steepest ascent, and in multivariable calculus, it is represented by the gradient vector of the elevation function.The gradient vector points...
Gradient Fields
A gradient field is a vector field derived from a scalar field. A scalar field assigns a single numerical value to every point in space, such as temperature, pressure, or electric potential. The gradient field describes how that value changes from point to point. It gives both the direction of the fastest increase and the rate of change in that direction.For a scalar field f(x, y), the gradient is written as\begin{equation*}\nabla f=\left\langle \jfrac{\partial f}{\partial x},\jfrac{\partial...
Maximizing the Directional Derivative
The directional derivative is a central concept in multivariable calculus that describes how a function changes at a given point when moving in a specified direction. This direction is represented by a unit vector, ensuring that only the orientation influences the rate of change. By varying the direction, different rates of change can be observed, demonstrating that the directional derivative depends strongly on the chosen direction.The directional derivative is computed using the gradient...
Significance of the Gradient Vector
A surface defined by a function of two variables can be understood by examining how it changes along specific directions. When one variable is held constant, the surface reduces to a curve that reflects variation in the other variable. For example, fixing one variable and moving parallel to a coordinate axis produces a cross-sectional curve. The slope of this curve at a given point represents how the function changes in that particular direction, providing a measure of local steepness.By...
Deconvolution
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Reducing Line Loss
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...