Related Experiment Video
Updated: Aug 21, 2026

Using the Open-Source MALDI TOF-MS IDBac Pipeline for Analysis of Microbial Protein and Specialized Metabolite Data
Published on: May 15, 2019
Dasinya: A dataset and preprocessing pipeline for the Badini Kurdish dialect
Vaman A Saeed1, Karwan Jacksi2
1Information Technology Department, Technical College of Duhok, Duhok Polytechnic University, Duhok, Kurdistan Region, Iraq.
None:
Despite notable advances in NLP for low-resource languages, dialects written in non-Latin scripts - especially those without standardised digital resources - continue to receive insufficient attention. The Badini Kurdish dialect, written in Perso-Arabic script and spoken across the Kurdistan Region of Iraq, lacks any dedicated NLP preprocessing pipeline in the literature. This paper presents Dasinya, the first openly released Badini Kurdish text corpus, together with the Badini Dialect Processing Toolkit (BDPT) - the first NLP preprocessing pipeline built specifically for this dialect. Dasinya contains 107 files, 87,545 sentences, and roughly 1.28 million words, spanning two source types and nine sub-genres: six book sub-genres (literary, historical, cultural, scientific, psychological, and educational) and three non-book sub-genres (news, academic/linguistic, and broadcast). These files were drawn from two repositories: the Badirxanian Public Library in Duhok (98 books) and the Semantic Web Lab at the University of Zakho (4 books and 5 non-book files). BDPT is structured as a six-stage pipeline that handles validation, Unicode normalisation - enforcing Badini-specific character mappings and preserving the Zero-Width Non-Joiner as a morpheme boundary marker - structural cleaning, deep cleaning, text preprocessing, and sentence segmentation. Orthographic problems particular to Badini Kurdish are addressed at each stage; none of these are handled by existing Arabic, Persian, or Sorani tools. At Stage 6, sentence segmentation reaches a specialist-validated Acc% of 91.4% across all 107 files, measured under a two-metric framework (OK% and Acc%) with genre-sensitive thresholds that reflect genuine linguistic differences between books, news, broadcast, and academic texts. Both the Dasinya corpus and the BDPT pipeline are openly available at https://doi.org/10.5281/zenodo.20729060 and are intended to underpin future Badini Kurdish NLP work, including tokenisation, part-of-speech tagging, stemming, and stopword removal.
