Related Experiment Video
Updated: Jun 5, 2025

05:33
Experimental Paradigm for Measuring the Effect of Induced Emotion on Grammar Learning
Published on: January 29, 2020
5.9K
An automated approach to identify sarcasm in low-resource language
Shumaila Khan1, Iqbal Qasim1, Wahab Khan1
1Institute of CS & IT, University of Science & Technology, Bannu, Pakistan.
Plos One
|December 5, 2024
Summary
This study explores machine learning for sarcasm detection in Urdu, a low-resource language. The Support Vector Machine (SVM) model achieved 0.85 accuracy, highlighting the need for language-specific approaches.
Area of Science:
- Natural Language Processing (NLP)
- Computational Linguistics
- Machine Learning (ML)
Background:
- Sarcasm detection research is predominantly in English, with limited exploration in low-resource languages like Urdu.
- Developing effective sarcasm detection for Urdu is challenging due to the scarcity of annotated datasets and unique linguistic nuances.
Purpose of the Study:
- To investigate the efficacy of various machine learning algorithms for sarcasm detection in the Urdu language.
- To address the data scarcity challenge by curating and releasing the Urdu Sarcastic Tweets (UST) Dataset.
- To establish a benchmark for future research in low-resource language sarcasm detection.
Main Methods:
- Curated and released the Urdu Sarcastic Tweets (UST) Dataset from user-generated social media comments.
- Evaluated baseline machine learning classifiers including Support Vector Machine (SVM), Decision Tree (DT), K-Nearest Neighbor (K-NN), Linear Regression (LR), Random Forest (RF), Naïve Bayes (NB), and XGBoost.
- Validated model performance on the newly created UST dataset and the Tanz-Indicator dataset.
Main Results:
- The Support Vector Machine (SVM) classifier demonstrated superior performance, achieving an accuracy of 0.85.
- SVM consistently outperformed other evaluated machine learning models across different experimental configurations.
- The study confirms the feasibility of applying ML techniques to Urdu sarcasm detection.
Conclusions:
- Machine learning models, particularly SVM, are effective for sarcasm detection in Urdu.
- Tailored approaches considering specific linguistic characteristics are crucial for low-resource languages.
- The release of the UST dataset facilitates further research and development in this domain.
Related Concept Videos
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Nonconscious Mimicry
4.5K
Nonconscious mimicry occurs when individuals alter their mannerisms to match the behaviors and expressions of those nearby, without intention.
4.5K
Nonsense-mediated mRNA Decay
10.5K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
10.5K

