Related Experiment Video
Updated: Apr 5, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.7K
Multi-Spectral Fusion Based Approach for Arbitrarily Oriented Scene Text Detection in Video Images
Summary
This study introduces a novel method for scene text detection, enhancing low-resolution text pixels using frequency domain analysis and wavelet sub-bands. The approach improves text detection accuracy in complex natural scenes and videos.
Area of Science:
- Computer Vision
- Image Processing
- Pattern Recognition
Background:
- Scene text detection is complex due to variations in background, contrast, font, and arbitrary text orientations.
- Multi-script text and low-resolution pixels further challenge existing detection methods.
Purpose of the Study:
- To develop an improved method for robust scene text detection in natural images and videos.
- To enhance the detection of low-resolution text pixels and handle complex text variations.
Main Methods:
- Convolving Laplacian with wavelet sub-bands in the frequency domain to enhance low-resolution text.
- Fusing spectral results from different sub-bands for candidate text pixel detection.
- Utilizing Maximally Stable Extremal Regions (MSER) and Stroke Width Transform (SWT) for candidate text region detection.
- Employing symmetry-driven nearest neighbor for text line restoration and alignment.
Main Results:
- The proposed method demonstrates superior performance compared to state-of-the-art techniques.
- Effective enhancement of low-resolution text pixels leading to improved detection rates.
- Successful text detection and line restoration on diverse datasets including ICDAR 2011, 2013, and MSRA-TD500.
Conclusions:
- The developed approach offers a significant advancement in scene text detection accuracy and robustness.
- The combination of frequency domain enhancement and region-based detection proves effective for complex text scenarios.