Related Experiment Video
Updated: Jun 13, 2025

06:31
Author Spotlight: Enhancing Rheumatoid Arthritis Research Through HR-pQCT Imaging Analysis
Published on: October 6, 2023
2.1K
GHCR-A dataset for Grantha handwritten character recognition.
Basaraboyina Yohoshiva1, Nagendra Panini Challa1
1VIT-AP University, Amaravati, Andhra Pradesh, India.
Data in Brief
|September 10, 2024
Summary
A new dataset of handwritten Grantha characters, including numbers and vowels, was created to aid machine learning research. This resource addresses a gap in available data for the Grantha script, supporting Indian language technology development.
Area of Science:
- Computer Science
- Linguistics
- Digital Humanities
Background:
- The Grantha script, historically significant in South Indian languages, lacks sufficient digital datasets for research.
- Existing resources for Grantha character recognition are limited, hindering advancements in related AI applications.
Purpose of the Study:
- To introduce a comprehensive dataset of handwritten Grantha characters (numbers and vowels).
- To provide a benchmark resource for developing and evaluating machine learning models for Grantha character recognition.
- To facilitate research in Indian languages connected to the Grantha script.
Main Methods:
- Collected handwritten samples of 10 Grantha numbers and 34 vowels from diverse age groups.
- Digitized and preprocessed 5852 images, including segmentation, resizing, and grayscale conversion.
- Organized data into image and CSV formats with corresponding labels for machine learning.
Main Results:
- The final dataset contains 1330 samples for numbers and 4522 samples for vowels.
- The dataset comprises 44 distinct Grantha characters, with 133 handwritten samples per character.
- Data is available in both image and CSV formats, suitable for immediate use.
Conclusions:
- This dataset fills a critical gap for Grantha script research, particularly in numeral and vowel recognition.
- It serves as a foundational resource for machine learning initiatives focused on Grantha-influenced Indian languages.
- The availability of this dataset is expected to spur further innovation in script recognition and natural language processing.

