Related Experiment Video
Updated: Jun 3, 2026

09:43
Transmission of Multiple Signals through an Optical Fiber Using Wavefront Shaping
Published on: March 20, 2017
ConvShareViT: A Vision Transformer-Like Architecture for Free-Space Optical Accelerators
Summary
This study introduces convolutional shared vision transformers (ConvShareViT), a novel deep learning model for optical systems. ConvShareViT achieves comparable attention scores to standard vision transformers (ViTs) and offers significantly faster inference speeds.
Area of Science:
- Computer Science
- Optical Engineering
- Artificial Intelligence
Background:
- Vision Transformer (ViT) architectures are powerful for image recognition but computationally intensive.
- Adapting deep learning models to free-space optical systems presents unique challenges.
- Existing methods may not fully leverage the potential of optical computing for AI tasks.
Purpose of the Study:
- To introduce a novel deep learning architecture, convolutional shared vision transformers (ConvShareViT), tailored for 4f free-space optical systems.
- To investigate the effectiveness of replacing linear layers in ViT with depthwise convolutional layers with shared weights.
- To evaluate the attention learning capabilities and inference speed of ConvShareViT compared to standard ViTs.
Main Methods:
- Developed ConvShareViT by replacing Multi-Head Self-Attention (MHSA) and Multilayer Perceptron (MLP) linear layers with depthwise convolutional layers.
- Employed shared weights across input channels in the convolutional layers.
- Conducted 12 experiments to systematically analyze the attention mechanism's effectiveness with different configurations, including valid-padded and same-padded convolutions.
- Compared quantitative attention scores and theoretical inference speeds against standard ViTs.
Main Results:
- ConvShareViT configurations utilizing valid-padded shared convolutions successfully learned attention mechanisms.
- Attention scores achieved by ConvShareViT were comparable to those of standard ViTs.
- Same-padded convolutions demonstrated limitations in attention learning, behaving more like traditional Convolutional Neural Networks (CNNs).
- ConvShareViT theoretically offers up to 3.04x faster inference compared to GPU-based systems.
Conclusions:
- ConvShareViT is a viable adaptation of the ViT architecture for optical deep learning applications.
- Shared convolutional layers, particularly with valid padding, can effectively implement attention mechanisms in optical systems.
- The proposed architecture presents a significant speed advantage, making it promising for future optical AI hardware.
Related Concept Videos
Acceleration Vectors
In everyday conversation, accelerating means speeding up. Acceleration is a vector in the same direction as the change in velocity, Δv, therefore the greater the acceleration, the greater the change in velocity over a given time. Since velocity is a vector, it can change in magnitude, direction, or both. Thus acceleration is a change in speed or direction, or both. For example, if a runner traveling at 10 km/h due east slows to a stop, reverses direction, and continues their run at 10 km/h due...
Confocal Fluorescence Microscopy
Confocal microscopy is an advanced microscopic technique. The prime advantage of the confocal microscope over other microscopy techniques is its ability to block the out-of-focus light from the illuminated samples using pinholes. It is widely used with fluorescence optics to obtain high-resolution, sharp contrast images. Unlike optical microscopes, confocal microscopes use a focused beam of light laser to scan the entire sample surface at different z-planes. These microscopes are, therefore,...
X-ray Imaging
German physicist Wilhelm Röntgen (1845–1923) was experimenting with electrical current when he discovered that a mysterious and invisible "ray" would pass through his flesh but leave an outline of his bones on a screen coated with a metal compound. In 1895, Röntgen made the first durable record of the internal parts of a living human: an "X-ray" image (as it came to be called) of his wife’s hand. Scientists worldwide quickly began their own experiments with X-rays, and by 1900, X-ray was widely...
Vision
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
Imaging Biological Samples with Optical Microscopy
Optical microscopy uses optic principles to provide detailed images of samples. Antonie van Leeuwenhoek designed the first compound optical microscope in the 17th century to visualize blood cells, bacteria, and yeast cells. In 1830, Joseph Jackson Lister created an essentially modern light microscope. The 20th century saw the development of microscopes with enhanced magnification and resolution.
In optical microscopy, the specimen to be viewed is placed on a glass slide and clipped on the stage...
In optical microscopy, the specimen to be viewed is placed on a glass slide and clipped on the stage...
Accelerators
Accelerators in concrete serve as admixtures to speed up the hardening process, enabling the concrete to achieve early strength faster. Although accelerators do not necessarily impact the time it takes concrete to set, they reduce this time in practice. A common accelerator is calcium chloride, which is particularly useful for hastening early strength development in cold weather or for rapid repair jobs that require quick heat generation after mixing.
The effectiveness of calcium chloride can...
The effectiveness of calcium chloride can...

