Related Experiment Video
Updated: May 7, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
A 48-class handwritten dataset for the endangered chakma language
Jannatul Ferdeous1, Md Abdul Kayum1, Ahmed Islam1
1Department of Computer Science and Engineering, Green University of Bangladesh, Purbachal American City, Kanchon 1460, Dhaka, Bangladesh.
Abstract:
This article describes a comprehensive handwritten character dataset for the endangered Chakma language, primarily spoken in the Chittagong Hill Tracts of Bangladesh. The dataset comprises 37,708 processed RGB images, standardized to 40 × 40 pixels, covering 48 distinct classes that include 38 letters and 10 numerals. Data collection involved manual input from a diverse demographic of native speakers in the Rangamati and Khagrachhari districts, utilizing standardized form to capture authentic stylistic variability. The raw handwritten samples were subsequently digitized using high-resolution scanning, followed by automated cropping and resizing scripts to generate a uniform, machine-learning-ready format. This open-access resource addresses the scarcity of digital tools for indigenous scripts and can be utilized for training Handwritten Character Recognition (HCR) models, developing synthetic fonts via generative networks, and conducting linguistic analysis of Chakma handwriting characteristics.
Related Concept Videos
Components of Language
Conservation of Small Populations
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Genetic Lingo
Conservation of Declining Populations
