Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

How Data are Classified: Categorical Data01:11

How Data are Classified: Categorical Data

47.6K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
47.6K
Tagging and Fusion Proteins01:24

Tagging and Fusion Proteins

8.7K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
8.7K
Peptide Identification Using Tandem Mass Spectrometry01:33

Peptide Identification Using Tandem Mass Spectrometry

8.8K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
8.8K
Data Collection I01:30

Data Collection I

8.9K
Data collection gathers information needed to make accurate judgments about a patient's present condition. During a health history interview, subjective data is collected from the patient, their caregivers, or family members, and objective data is collected through observations and physical assessment. Patients are the primary source of subjective data. Thus information gathered from patients through interviews, observations, and physical examination is primary data. Secondary sources of...
8.9K
Classification of Signals01:30

Classification of Signals

1.5K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.5K
Aggregates Classification01:29

Aggregates Classification

1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Dataset of sentiment tagged language resources for Macedonian language.

Data in brief·2026
Same author

Dataset of Uzbek verbs with formation and suffixes.

Data in brief·2025
Same author

Dataset of vocabulary in Uzbek primary education: Extraction and analysis in case of the school corpus.

Data in brief·2025
Same author

Dataset of Edmonds' bi-vectors and tri-vectors with realizations.

Data in brief·2024
Same author

Dataset of sentiment tagged language resources for Bosnian language.

Data in brief·2024
Same author

Parallel texts dataset for Uzbek-Kazakh machine translation.

Data in brief·2024

Related Experiment Video

Updated: Jul 2, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

UzbekPOS: A multi-domain dataset for Uzbek part-of-speech tagging.

Maksud Sharipov1, Elmurod Kuriyozov1, Jernej Vičič2

  • 1Urgench State University named after Abu Rayhan Biruni, 14, Kh. Alimdjan str, Urgench 220100, Uzbekistan.

Data in Brief
|March 19, 2026
PubMed
Summary

This paper introduces UzbekPOS, the largest publicly available part-of-speech (POS) tagged dataset for the Uzbek language. This resource aids natural language processing and linguistic studies for morphologically complex Turkic languages.

Keywords:
Morphological annotationNatural language processingPOS taggingUzbek language

More Related Videos

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

Mining Spatial Transcriptomics Datasets using DeepSpaceDB
10:16

Mining Spatial Transcriptomics Datasets using DeepSpaceDB

Published on: September 5, 2025

Related Experiment Videos

Last Updated: Jul 2, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

Mining Spatial Transcriptomics Datasets using DeepSpaceDB
10:16

Mining Spatial Transcriptomics Datasets using DeepSpaceDB

Published on: September 5, 2025

Area of Science:

  • Computational Linguistics
  • Corpus Linguistics
  • Artificial Intelligence

Background:

  • Uzbek, a morphologically complex Turkic language, is under-resourced in NLP.
  • Existing Uzbek language resources lack comprehensive part-of-speech (POS) tagged data.

Purpose of the Study:

  • Introduce UzbekPOS, the largest publicly available POS-tagged dataset for Uzbek.
  • Provide a foundational resource for NLP, AI, and corpus linguistics applications for Uzbek.
  • Facilitate POS tagging, morphological analysis, and cross-linguistic studies in Turkic languages.

Main Methods:

  • Manual annotation of sentences from diverse Uzbek text sources by professional annotators.
  • Development of a finely-grained POS tagset integrating Universal Dependencies with Uzbek-specific labels (16 tags total).
  • Cross-verification of annotations by at least two annotators to ensure high reliability.

Main Results:

  • The UzbekPOS dataset contains nearly 4.5K sentences and over 53K token/tag pairs.
  • The dataset is available in multiple formats (txt, TSV, JSON, conllu) for broad usability.
  • This represents one of the first and largest openly published POS-tagged corpora for Uzbek.

Conclusions:

  • UzbekPOS serves as a crucial resource for training POS taggers and evaluating machine learning models for Uzbek.
  • The dataset's reusability extends to morphological analysis, syntactic parsing, and transfer learning within Turkic languages.
  • UzbekPOS can act as seed material for creating similar corpora for other Turkic languages, fostering cross-linguistic research.