Related Experiment Video
Updated: Jul 26, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Identification of offensive language in Urdu using semantic and embedding models.
Sajid Hussain1, Muhammad Shahid Iqbal Malik1, Nayyer Masood1
1Department of Computer Science, Capital university of Science and Technology, Islamabad, Pakistan.
This study introduces a new dataset for detecting offensive Urdu language, achieving 90% accuracy with a hybrid model combining TF-IDF, bag-of-words, and word2vec features. This advancement offers practical applications for online safety and future research.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Social Media Analysis
Background:
- Automatic detection of offensive language is crucial for online safety.
- Generalizing solutions across languages is challenging due to linguistic diversity.
- Prior research on offensive language detection has primarily focused on Western languages, with limited work on low-resource languages like Urdu.
Purpose of the Study:
- To develop and evaluate a robust system for offensive language detection in Urdu.
- To create a new, comprehensive dataset for Urdu offensive language identification.
- To explore and compare various feature engineering techniques for this task.
Main Methods:
- A new dataset of 7,500 Urdu Facebook posts was curated for offensive language detection.
- Four feature engineering models were employed: Term Frequency-Inverse Document Frequency (TF-IDF), Bag-of-Words, Word n-grams, and Word2Vec embeddings.
- Experiments utilized standalone models and ensemble methods, including stacking and wrapper-based feature selection.
Main Results:
- The Word2Vec embedding model, when used with a stacking ensemble, achieved 88.27% accuracy as a standalone model.
- A hybrid model combining TF-IDF, Bag-of-Words, and Word2Vec features reached 90% accuracy and 97% Area Under the Curve (AUC).
- The proposed hybrid approach significantly outperformed baseline methods across multiple performance metrics.
Conclusions:
- The developed hybrid feature model demonstrates superior performance for Urdu offensive language detection.
- The creation of a dedicated Urdu dataset addresses a gap in low-resource language NLP research.
- Findings have practical implications for developing content moderation tools and enhancing online safety in Urdu-speaking communities.
More Related Videos
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Stereotype Content Model
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Components of Language
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...