Related Experiment Video
Updated: Aug 2, 2025

Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
Sequence-to-sequence pretraining for a less-resourced Slovenian language
Matej Ulčar1, Marko Robnik-Šikonja1
1Faculty of Computer and Information Science, University of Ljubljana, Ljubljana, Slovenia.
New Slovene T5 (SloT5) models show promise for text generation tasks despite challenges with limited data. These sequence-to-sequence models offer valuable results for Slovene natural language processing, especially for generative applications.
Area of Science:
- Natural Language Processing
- Machine Learning
- Computational Linguistics
Background:
- Large pretrained language models like BERT and T5 have advanced NLP.
- T5's sequence-to-sequence objective is well-suited for text generation.
- Existing T5 models are often limited to well-resourced languages.
Purpose of the Study:
- To develop and evaluate T5-type sequence-to-sequence models for the Slovene language.
- To assess the performance of these models on various classification and generation tasks.
- To compare SloT5 models against multilingual and monolingual baselines.
Main Methods:
- Trained two sizes of T5-type models for Slovene (SloT5).
- Evaluated models on 11 tasks: 8 classification (NER, sentiment, QA, NLI, coreference, lemmatization) and 3 generation (simplification, summarization).
- Compared SloT5 against mT5, mBART-50, multilingual BERT, XLM-RoBERTa, a trilingual BERT, and monolingual SloBERTa.
Main Results:
- SloT5 models generally underperformed the monolingual SloBERTa on classification tasks.
- SloT5 models demonstrated effectiveness in text generation tasks, yielding useful results.
- Model performance is influenced by size, and insufficient Slovene training data limits large model pretraining.
Conclusions:
- SloT5 models are valuable for Slovene text generation, despite limitations in classification.
- The findings suggest potential generalizability to other low-resource languages.
- Training code and models are publicly released to facilitate further research.
More Related Videos
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
05:48Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Related Concept Videos
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Purposive Learning
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Initiation of Translation