Related Experiment Video
Updated: Jul 24, 2025

Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning
Published on: January 29, 2020
Corpus creation and language identification for code-mixed Indonesian-Javanese-English Tweets.
Ahmad Fathan Hidayatullah1,2, Rosyzie Anna Apong1, Daphne T C Lai1
1School of Digital Science, Universiti Brunei Darussalam, Bandar Seri Begawan, Brunei.
This study introduces a new model for language identification in social media text, specifically for code-mixed Indonesian, Javanese, and English. Fine-tuned IndoBERTweet models show superior performance in identifying languages within these mixed-language tweets.
Area of Science:
- Computational Linguistics
- Natural Language Processing (NLP)
- Social Media Analysis
Background:
- Code-mixing, the phenomenon of language mixing, is increasingly prevalent in social media text.
- This linguistic trend poses significant challenges for Natural Language Processing (NLP) tasks, particularly Language Identification (LID).
- Existing LID models often struggle with the complexities of code-mixed data, necessitating specialized approaches.
Purpose of the Study:
- To develop and evaluate a word-level language identification model for code-mixed Indonesian, Javanese, and English tweets.
- To introduce a novel annotated corpus, the Indonesian-Javanese-English Language Identification (IJELID) dataset, for training and evaluation.
- To compare the effectiveness of different modeling strategies, including fine-tuned BERT, BLSTM, and CRF, for this specific task.
Main Methods:
- Creation of the Indonesian-Javanese-English Language Identification (IJELID) corpus, detailing data collection and annotation standards.
- Investigation of various machine learning approaches for word-level language identification.
- Fine-tuning of BERT-based models (IndoBERTweet), Bidirectional Long Short-Term Memory (BLSTM) networks, and Conditional Random Fields (CRF).
Main Results:
- Fine-tuned IndoBERTweet models demonstrated superior performance in identifying languages within code-mixed Indonesian, Javanese, and English tweets compared to BLSTM and CRF.
- The effectiveness of BERT models is attributed to their ability to capture contextual information of each word within a text sequence.
- Sub-word language representation within BERT models proved to be a reliable mechanism for language identification in code-mixed text.
Conclusions:
- Fine-tuned IndoBERTweet models offer a robust solution for language identification in code-mixed social media text.
- The developed IJELID corpus provides a valuable resource for future research in code-mixed language identification.
- BERT's contextual understanding and sub-word representation capabilities are key to successfully addressing the challenges of code-mixing in NLP.
More Related Videos
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Language and Cognition
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Genetic Lingo