Related Experiment Videos
A multimodal learning framework for Arabic handwritten word recognition and future research directions
1Department of Computer Engineering, College of Computer Engineering and Science, Prince Mohammad Bin Fahd University, Khobar, Saudi Arabia. jalkhateeb@pmu.edu.sa.
Abstract:
In fact, Arabic handwritten word recognition remains a challenging task due to the cursive nature of the Arabic script, positional character variations, diacritical marks, and significant intra- and inter-writer variability. While both probabilistic models and deep learning approaches have been widely explored, their complementary strengths and integration into scalable recognition systems remain insufficiently addressed. This paper presents a unified recognition framework that integrates a classical Hidden Markov Model (HMM)-based pipeline with a deep Convolutional Neural Network (CNN)-based visual learning approach for offline Arabic handwritten word recognition. The framework is evaluated under identical experimental conditions using a real-world dataset of 5000 handwritten word images of Libyan city names, collected from writers with diverse backgrounds. The experimental results show that the CNN-based model achieves a recognition accuracy of 95.4%, whereas the HMM-based system achieves an accuracy of 83.62% indicating the significant efficacy of learned visual representations within the assessed experimental scenario. The paper examines the advantages and disadvantages of data-driven feature learning and probabilistic sequence modeling in addition to performance benchmarking. A research roadmap pointing out sequence-aware deep architectures, open-vocabulary recognition, and language model integration toward scalable handwritten word recognition systems is carried out as a consequence of these findings.