Related Experiment Video
Updated: Jan 21, 2026

An Experimental Paradigm for Measuring the Effects of Ageing on Sentence Processing
Published on: October 25, 2019
Low-resource Language Identification of English and Kokborok Code-Mixed Sentences
1Department of Computer Science & Engineering, National Institute of Technology; enjula.phd23@nitap.ac.in.
Abstract:
There is a growing need for a model for automatic word-level language detection due to the growing usage of multilingual text. To identify code-switched and code-mixed sentences in the Kokborok language and English at the word level, we propose an unsupervised model. Numerous studies on code-mixed text have been published, including for several Indian languages. However, to the best of our knowledge, our work represents a pioneering effort, being the first to identify languages for low-resource English-Kokborok language pairs. We have employed a method that merges data from a frequency dictionary with data from a character n-gram model. The Viterbi technique combines a character n-gram Markov model with a frequency lexicon to aid in precise language identification at the word level. The proposed method performed well, with a word-level accuracy of 93.15%, which is better than the BiLSTM-CRF model of 88.5% and the rule-based baseline of 85.3%. This demonstrates the effectiveness of combining lexicon and statistical methodology for handling low-resource languages without depending on massive annotated datasets.
More Related Videos
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Short-distance Transport of Resources
lncRNA - Long Non-coding RNAs
lncRNA - Long Non-coding RNAs
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

