Related Experiment Video
Updated: Dec 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Enhancing African low-resource languages: Swahili data for language modelling.
Casper S Shikali1,2,3, Refuoe Mokhosi1
1School of information and Software Engineering, University of Electronic Science and Technology of China., Xiyuan Ave, West Hi-Tech Zone, 611731 Chengdu, Sichuan, PR China.
This study introduces new Swahili datasets for natural language processing (NLP), addressing the lack of resources for low-resource languages. These datasets aim to improve Swahili language models and other NLP tasks.
Area of Science:
- Computational Linguistics
- African Languages
- Natural Language Processing
Background:
- Neural network-based language modeling requires substantial data for effective word representation in Natural Language Processing (NLP).
- African languages, like Swahili, are often classified as low-resource languages due to insufficient data for NLP tasks.
- Existing NLP resources are scarce for Swahili, hindering its development and application.
Purpose of the Study:
- To address the data scarcity for Swahili in NLP by creating and contributing new datasets.
- To provide essential resources for improving Swahili language models and other NLP applications.
- To support research and development for low-resource languages in the field of NLP.
Main Methods:
- An unannotated Swahili dataset was derived through the pre-processing of raw Swahili text using a Python script.
- A Swahili syllabic alphabet was formulated.
- A Swahili word analogy dataset was developed, drawing inspiration from an existing English dataset.
Main Results:
- The creation and contribution of three novel Swahili datasets: an unannotated dataset, a syllabic alphabet, and a word analogy dataset.
- Demonstrated a method for generating NLP resources for low-resource languages.
- Provided foundational data for advancing Swahili NLP.
Conclusions:
- The newly derived Swahili datasets are expected to significantly benefit Swahili language models.
- These resources will also support various downstream NLP tasks, including part-of-speech tagging, machine translation, and sentiment analysis.
- This work contributes to bridging the resource gap for low-resource African languages in NLP.
More Related Videos
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Improving Translational Accuracy
Improving Translational Accuracy
Language and Cognition
Components of Language
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...

