Related Experiment Videos
Enhancing camera-captured Devanagari documents via geometric filtering for improved vision-language model text
Anup Kelkar1, Parag Deshpande2, O G Kakde2
1Department of Computer Science and Engineering, Indian Institute of Information Technology (IIIT), Nagpur, Maharashtra, 440024, India.
Abstract:
Despite significant advances, modern Vision-Language Model (VLM) platforms remain constrained by the scarcity of training data in non-English languages, limiting their global applicability in Text extraction for regional languages written in Devanagari scripts. Recognizing Devanagari text, particularly in literary materials, poses significant challenges in achieving high accuracy and efficient execution. Effective Text extraction in current VLM platforms must perform several pre-processing steps to retain crucial information while preserving sentence structure and meaning. • This study proposes an image enhancement method for Devanagari text based on geometric filters and clustering algorithms. • By leveraging the unique geometrical characteristics of Devanagari characters, the technique accurately detects Devanagari sentences in various font sizes, orientations, and degraded hand-held camera-captured images. It is robust to variations in document rotation, size, and color. • The method was tested on a dataset of images from books, papers, brochures, and pamphlets, demonstrating superior performance over existing OCR pre-processing techniques for current VLM platforms. The proposed approach achieved a success rate of 95.23 % in locating Devanagari sentences, improving the accuracy of current VLM platforms Text extraction.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy