Related Experiment Video
Updated: Sep 6, 2026

Transcranial Direct Current Stimulation (tDCS) of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
A POS-tagged dataset of Uzbek word lemmas and stems for morphological and NLP research
1Urgench State University named after Abu Rayhan Biruni, Khamid Alimdjan, 14, 220100, Urgench City, Uzbekistan.
Abstract:
This article presents a comprehensive, multi-domain dataset of stems and lemmas for the Uzbek language, systematically categorized by parts-of-speech (POS). The dataset contains 46,307 unique stems and 64,052 corresponding lemmas classified into 12 distinct grammatical categories, including independent, auxiliary, and intermediate word classes. Data were methodically extracted from authoritative sources such as the Explanatory Dictionary of the Uzbek Language and various digital text corpora. Following a rigorous right-to-left suffix stripping methodology, words were segmented to isolate the base stems, and their corresponding canonical dictionary forms (lemmas) were identified. The finalized linguistic data is structured in a machine-readable tabular (.xlsx) format consisting of three essential columns: stem, lemmas, and POS tag. This dataset provides a foundational open-source resource for natural language processing (NLP) applications in the morphologically complex, agglutinative Uzbek language, actively supporting the development, training, and benchmarking of rule-based and machine-learning tools such as lemmatizers, stemmers, and POS taggers.
