Related Experiment Video
Updated: Aug 31, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
DravidianCodeMix: sentiment analysis and offensive language identification dataset for Dravidian languages in
Bharathi Raja Chakravarthi1, Ruba Priyadharshini2, Vigneshwaran Muralidaran3
1Insight SFI Research Centre for Data Analytics, Data Science Institute, National University of Ireland Galway, Galway, Ireland.
Researchers created a new dataset of over 60,000 social media comments in Tamil, Kannada, and Malayalam for sentiment analysis and offensive language identification. This resource aids natural language processing for under-resourced Dravidian languages.
Area of Science:
- Computational Linguistics
- Natural Language Processing
- Social Media Analysis
Background:
- Under-resourced Dravidian languages present significant challenges for NLP tasks.
- Social media platforms are rich sources of multilingual user-generated content, often exhibiting code-mixing.
- Existing datasets often lack sufficient coverage for low-resource languages and specific annotation tasks.
Purpose of the Study:
- To develop and release a novel, manually annotated dataset for sentiment analysis and offensive language identification.
- To support research in natural language processing for Tamil, Kannada, and Malayalam.
- To capture code-mixing phenomena prevalent in multilingual social media data.
Main Methods:
- Manual annotation of over 60,000 YouTube comments by volunteer annotators.
- Annotation focused on sentiment analysis and offensive language identification.
- Dataset creation involved collecting and processing comments in Tamil-English, Kannada-English, and Malayalam-English.
Main Results:
- A comprehensive multilingual dataset comprising 44,000 Tamil-English, 7,000 Kannada-English, and 20,000 Malayalam-English comments.
- High inter-annotator agreement (Krippendorff's alpha) demonstrating annotation quality.
- Baseline machine learning and deep learning experiments established performance benchmarks.
Conclusions:
- The released dataset provides a valuable resource for advancing NLP research in under-resourced Dravidian languages.
- The dataset's inclusion of code-mixing phenomena offers unique opportunities for studying language interaction.
- Availability on GitHub and Zenodo promotes accessibility and further research.
Related Concept Videos
Genetic Lingo
LTR Retrotransposons
The internal coding region of LTR retrotransposons and their mechanism of transposition closely resembles a...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Non-LTR Retrotransposons
Language and Cognition

