Related Experiment Video
Updated: Sep 25, 2025

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Social media text analytics of Malayalam-English code-mixed using deep learning
S Thara1, Prabaharan Poornachandran1
1Department of Computer Science and Engineering, Amrita Vishwa Vidyapeetham, Amritapuri, India.
This study introduces a novel framework for analyzing Malayalam-English code-mixed social media text, achieving high accuracy in offensive language identification and sentiment analysis. The approach enhances understanding of informal digital communication.
Area of Science:
- Computational Linguistics
- Natural Language Processing
- Social Media Analysis
Background:
- Social media text, often code-mixed, presents significant challenges for automated analysis due to informal language and vocabulary.
- Existing methods struggle with the nuances of Malayalam-English code-mixed data, necessitating specialized approaches for tasks like offensive language identification and sentiment analysis.
Purpose of the Study:
- To develop and evaluate a framework for offensive language identification and sentiment analysis on Malayalam-English code-mixed text.
- To investigate the impact of embedding methods, deep learning architectures, and translation techniques on model performance.
Main Methods:
- Utilized Word2Vec and FastText for feature engineering, exploring dependencies among embeddings.
- Compared various deep learning models, including uni-directional, bi-directional, hybrid, and transformer approaches.
- Incorporated selective translation and transliteration, alongside hyper-parameter optimization.
Main Results:
- Achieved high F1-Scores: 0.76 for the Forum for Information Retrieval Evaluation (FIRE) 2020 dataset and 0.99 for the European Chapter of the Association for Computational Linguistics (EACL) 2021 dataset.
- The proposed strategy outperformed benchmarked models for Malayalam-English code-mixed messages.
- Detailed error analysis provided insights into model limitations and areas for improvement.
Conclusions:
- The developed framework demonstrates significant effectiveness in analyzing Malayalam-English code-mixed text for critical NLP tasks.
- This work represents a substantial advancement in processing informal, multilingual social media data, contributing to societal well-being.
- The findings highlight the importance of tailored methods for handling code-mixed language in digital communication.
More Related Videos
08:53Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Related Concept Videos
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...