Related Experiment Video
Updated: May 22, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Localizing text in scene images by boundary clustering, stroke segmentation, and string fragment classification
1Graduate Center, City University of New York, New York, NY 10016, USA. cyi@gc.cuny.edu
Summary
This study introduces a new framework for text extraction from complex images. The method improves text localization accuracy across diverse image types, outperforming existing algorithms.
Area of Science:
- Computer Vision
- Image Processing
- Pattern Recognition
Background:
- Extracting text from complex scene images is challenging due to varied backgrounds and text appearances.
- Existing text localization methods often struggle with intricate visual data.
Purpose of the Study:
- To propose a novel framework for robust text region extraction from complex scene images.
- To enhance text localization performance across diverse image datasets.
Main Methods:
- A three-step framework: boundary clustering (BC), stroke segmentation, and string fragment classification.
- Utilizes a bigram-color-uniformity method for boundary clustering.
- Employs Gabor-based text features derived from gradient, stroke distribution, and width for classification.
Main Results:
- The proposed framework effectively extracts text regions from images with complex backgrounds.
- Demonstrated superior performance compared to state-of-the-art text localization algorithms.
- Validated across scene images, born-digital images, broadcast video, and images captured by visually impaired individuals.
Conclusions:
- The novel framework offers a significant advancement in text extraction from challenging image sources.
- The approach shows promise for applications requiring accurate text localization in real-world scenarios.
- The method's robustness is confirmed by its performance on diverse and difficult datasets.
