Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Language Development01:22

Language Development

713
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
713
Improving Translational Accuracy02:07

Improving Translational Accuracy

13.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
13.7K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.4K
3.4K
Language and Cognition01:27

Language and Cognition

612
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
612
Components of Language01:24

Components of Language

655
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
655
Language01:16

Language

795
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
795

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Sesotho news headlines dataset for sentiment analysis.

Data in brief·2024
See all related articles

Related Experiment Video

Updated: Dec 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

895

Enhancing African low-resource languages: Swahili data for language modelling.

Casper S Shikali1,2,3, Refuoe Mokhosi1

  • 1School of information and Software Engineering, University of Electronic Science and Technology of China., Xiyuan Ave, West Hi-Tech Zone, 611731 Chengdu, Sichuan, PR China.

Data in Brief
|July 17, 2020
PubMed
Summary

This study introduces new Swahili datasets for natural language processing (NLP), addressing the lack of resources for low-resource languages. These datasets aim to improve Swahili language models and other NLP tasks.

Keywords:
Deep learningLanguage modellingNatural language processingNeural networksSyllablesUnannotated dataWord analogy

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

723
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

3.0K

Related Experiment Videos

Last Updated: Dec 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

895
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

723
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

3.0K

Area of Science:

  • Computational Linguistics
  • African Languages
  • Natural Language Processing

Background:

  • Neural network-based language modeling requires substantial data for effective word representation in Natural Language Processing (NLP).
  • African languages, like Swahili, are often classified as low-resource languages due to insufficient data for NLP tasks.
  • Existing NLP resources are scarce for Swahili, hindering its development and application.

Purpose of the Study:

  • To address the data scarcity for Swahili in NLP by creating and contributing new datasets.
  • To provide essential resources for improving Swahili language models and other NLP applications.
  • To support research and development for low-resource languages in the field of NLP.

Main Methods:

  • An unannotated Swahili dataset was derived through the pre-processing of raw Swahili text using a Python script.
  • A Swahili syllabic alphabet was formulated.
  • A Swahili word analogy dataset was developed, drawing inspiration from an existing English dataset.

Main Results:

  • The creation and contribution of three novel Swahili datasets: an unannotated dataset, a syllabic alphabet, and a word analogy dataset.
  • Demonstrated a method for generating NLP resources for low-resource languages.
  • Provided foundational data for advancing Swahili NLP.

Conclusions:

  • The newly derived Swahili datasets are expected to significantly benefit Swahili language models.
  • These resources will also support various downstream NLP tasks, including part-of-speech tagging, machine translation, and sentiment analysis.
  • This work contributes to bridging the resource gap for low-resource African languages in NLP.