Related Experiment Video
Updated: Jul 19, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
1.9K
Unsupervised Low-Light Video Enhancement With Spatial-Temporal Co-Attention Transformer
Summary
This study introduces LightenFormer, the first unsupervised method for low-light video enhancement. It uses a spatial-temporal transformer to improve brightness and temporal consistency, overcoming limitations of supervised Convolution Neural Networks (CNNs).
Area of Science:
- Computer Vision
- Artificial Intelligence
- Image Processing
Background:
- Supervised Convolution Neural Networks (CNNs) dominate low-light video enhancement but struggle with real-world data due to synthetic training.
- Existing methods exhibit temporal inconsistency, like flickering and motion blur, especially with large motions, due to CNNs' limited perception of long-range dependencies.
Purpose of the Study:
- To develop the first unsupervised method for low-light video enhancement.
- To enhance video brightness while maintaining temporal consistency, even with large motions.
- To overcome the generalization limitations of supervised methods trained on synthetic data.
Main Methods:
- Proposed LightenFormer, an unsupervised low-light video enhancement method utilizing a spatial-temporal co-attention transformer.
- Introduced S-curve Estimation Network (SCENet) for adaptive dynamic range adjustment.
- Developed Spatial-Temporal Refinement Network (STRNet) with a novel Spatial-Temporal Co-attention Transformer (STCAT) for temporal consistency and long-range dependency modeling.
- Designed two non-reference loss functions for unsupervised training, leveraging S-curve invertibility and noise independence.
Main Results:
- LightenFormer effectively enhances brightness and maintains temporal consistency in low-light videos.
- The spatial-temporal co-attention transformer captures long-range spatial and temporal correlations for improved motion handling.
- Outperformed state-of-the-art methods on SDSD and LLIV-Phone datasets in extensive experiments.
Conclusions:
- LightenFormer offers a novel unsupervised approach to low-light video enhancement, addressing key limitations of prior supervised methods.
- The method demonstrates superior performance in brightness enhancement and temporal stability.
- Paves the way for more robust and generalizable low-light video enhancement techniques in real-world scenarios.
Related Concept Videos
Transformers
1.1K
A device that transforms voltages from one value to another using induction is called a transformer. A transformer consists of two separate coils, or windings, wrapped around the same soft iron core. However, they are electrically insulated from each other.
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
1.1K
Light Acquisition
8.5K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
8.5K
Deconvolution
188
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
188
Transformers with Off-Nominal Turns Ratios
176
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
176
Types Of Transformers
1.0K
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
1.0K
Depth Perception and Spatial Vision
720
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
720

