Related Experiment Videos
SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition
Yasser Alhabashi1, Omer Nacar2, Serry Sibaee1
1RIOTU Lab, Prince Sultan University, Riyadh, 11586, Saudi Arabia.
Abstract:
Arabic Optical Character Recognition (OCR) plays a critical role in converting extensive Arabic print media into digital formats. The development of modern OCR models, particularly advanced vision-language models, is hindered by the lack of large, diverse, and well-structured datasets that accurately replicate real-world book layouts. Most existing Arabic OCR datasets focus on isolated words or lines and are limited in scale, typographic diversity, or structural complexity representative of books. To address this limitation, this paper introduces SARD (Large-Scale Synthetic Arabic OCR Dataset). SARD is a synthetically generated dataset specifically designed to simulate book-style documents, comprising 2,621,075 document images and 794.6 million words rendered in six fonts: Sakkal Majalla, Arial, Calibri, Scheherazade New, Amiri, and Traditional Arabic. Unlike datasets derived from scanned documents, SARD is free from real-world noise and distortions, providing a clean and controlled environment for model training. Its synthetic nature enables exceptional scalability and precise control over layout and content variation. This study describes the dataset's composition, generation process, and technical validation, and presents benchmark results for several traditional and deep-learning OCR models to highlight the challenges and opportunities associated with the dataset. SARD serves as a valuable resource for developing and evaluating robust OCR and vision-language models that can process diverse Arabic book-style texts.