Related Experiment Video
Updated: May 6, 2026

05:58
Digital Handwriting Analysis of Characters in Chinese Patients with Mild Cognitive Impairment
Published on: March 11, 2021
4.5K
A scarce dataset for ancient Arabic handwritten text recognition.
Rayyan Najam1, Safiullah Faizullah1
1Department of Computer Science, Islamic University, Madinah 42351, Saudi Arabia.
Data in Brief
|September 10, 2024
Summary
This study introduces a novel dataset of ancient Arabic manuscripts for Optical Character Recognition (OCR) research. This resource aids in developing and validating deep learning models for Arabic OCR and text correction.
Area of Science:
- Computer Science
- Artificial Intelligence
- Digital Humanities
Background:
- Deep learning models are advancing Optical Character Recognition (OCR) capabilities.
- A significant gap exists in publicly available datasets for ancient Arabic manuscript OCR.
- Existing Arabic OCR research is hindered by the lack of specialized datasets.
Purpose of the Study:
- To address the scarcity of ancient Arabic manuscript data for OCR.
- To provide a valuable resource for training, validating, and testing deep learning models.
- To support research in Arabic OCR and Arabic text correction.
Main Methods:
- Collected and curated a dataset of ancient Arabic manuscripts.
- Acquired images and expert-transcribed textual ground truth.
- Dataset comprises eight ancient books totaling forty pages.
Main Results:
- Successfully created and provided a unique dataset of ancient Arabic manuscripts.
- The dataset includes both images and corresponding textual ground truth.
- The collection spans diverse geographies and historical periods.
Conclusions:
- The introduced dataset significantly contributes to the field of Arabic OCR.
- It enables researchers to develop, augment, test, and generalize deep learning models.
- This resource is crucial for advancing Arabic OCR and text correction technologies.

