Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

2.6K
2.6K
Translation01:31

Translation

14.9K
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Proteins are...
14.9K
Polytene Chromosomes02:04

Polytene Chromosomes

3.0K
3.0K
Crossover Experiments01:16

Crossover Experiments

2.8K
Crossover experiments, also called the repeated-measurements design, is a study design in which all experimental units are exposed to all treatments in different periods. Crossover experiments are generally used in psychology, the pharmaceutical industry, agriculture, and medicine.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
2.8K
Multiple Comparison Tests01:13

Multiple Comparison Tests

3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Test for Homogeneity01:23

Test for Homogeneity

2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Annotated universal dependencies dataset for literary and educational Uzbek texts.

Data in brief·2026
Same author

UzbekPOS: A multi-domain dataset for Uzbek part-of-speech tagging.

Data in brief·2026
See all related articles

Related Experiment Video

Updated: Jul 1, 2025

Author Spotlight: Polysome Profiling Protocol for Studying Translational Regulation in Arabidopsis Under Heat Stress
08:39

Author Spotlight: Polysome Profiling Protocol for Studying Translational Regulation in Arabidopsis Under Heat Stress

Published on: October 11, 2024

1.6K

Parallel texts dataset for Uzbek-Kazakh machine translation.

Bobur Allaberdiev1, Gayrat Matlatipov2, Elmurod Kuriyozov2,3

  • 1National University of Uzbekistan named after Mirzo Ulugbek, Universitet Street, 4, Olmazor district, 100174, Tashkent city, Uzbekistan.

Data in Brief
|March 1, 2024
PubMed
Summary

A new parallel corpus for Uzbek and Kazakh was created for machine translation. This valuable dataset, available on Hugging Face, aids research in low-resource Turkic languages.

Keywords:
Kazakh languageMachine translationNLPParallel corpusText datasetTurkic languagesUzbek language

More Related Videos

Analysis of Translation Initiation During Stress Conditions by Polysome Profiling
10:59

Analysis of Translation Initiation During Stress Conditions by Polysome Profiling

Published on: May 19, 2014

18.3K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

555

Related Experiment Videos

Last Updated: Jul 1, 2025

Author Spotlight: Polysome Profiling Protocol for Studying Translational Regulation in Arabidopsis Under Heat Stress
08:39

Author Spotlight: Polysome Profiling Protocol for Studying Translational Regulation in Arabidopsis Under Heat Stress

Published on: October 11, 2024

1.6K
Analysis of Translation Initiation During Stress Conditions by Polysome Profiling
10:59

Analysis of Translation Initiation During Stress Conditions by Polysome Profiling

Published on: May 19, 2014

18.3K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

555

Area of Science:

  • Computational Linguistics
  • Natural Language Processing
  • Machine Translation

Background:

  • Developing parallel corpora is crucial for advancing machine translation (MT) capabilities, especially for low-resource languages.
  • Existing resources for Uzbek and Kazakh are limited, hindering the development of effective MT systems for these Turkic languages.

Purpose of the Study:

  • To present a novel, high-quality parallel corpus for Uzbek and Kazakh raw texts.
  • To detail the methodology employed in constructing this dataset.
  • To highlight the corpus's potential for reuse in various Natural Language Processing (NLP) applications, particularly machine translation.

Main Methods:

  • A multi-stage data collection approach was utilized, combining existing parallel data, openly available resources (literature, web news), and expert manual translation.
  • Sentence alignment techniques were applied to integrate data from diverse sources.
  • The final corpus was curated and made accessible via the Hugging Face platform.

Main Results:

  • A substantial parallel corpus of Uzbek-Kazakh raw texts has been successfully created and published.
  • The dataset encompasses a wide range of topics and genres, increasing its utility.
  • The corpus is suitable for training and testing statistical and neural machine translation models.

Conclusions:

  • The developed Uzbek-Kazakh parallel corpus is a significant contribution to NLP research for low-resource Turkic languages.
  • Its availability on Hugging Face promotes accessibility and collaborative development.
  • This resource can serve as a foundational model for creating similar corpora for other related languages.