Related Experiment Video
Updated: Jan 17, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Cword2vec: a novel morphological rule-based word embedding approach for Urdu text sentiment analysis
Saquib Khushhal1, Abdul Majid1, Syed Ali Abass1
1Department of Computer Science and Information Technology, University of Azad Jammu & Kashmir, Pakistan, Muzaffarabad, Pakistan.
This study introduces Cword2vec, a novel method for Urdu natural language processing, improving sentiment analysis by effectively handling complex compound words. Morphological rule-based embeddings significantly outperform traditional techniques.
Area of Science:
- Natural Language Processing (NLP)
- Computational Linguistics
- Machine Learning
Background:
- Word embeddings are crucial for NLP, capturing syntactic and semantic word information.
- Urdu, spoken by over 231 million people, lacks sufficient NLP research, especially concerning its complex word structures.
- Urdu's unspecified word boundaries and prevalence of compound words pose challenges for traditional segmentation methods like bigrams and trigrams.
Purpose of the Study:
- To address the challenges of Urdu compound words in NLP.
- To propose a novel morphological rule-based compound word embedding (Cword2vec) for Urdu text representation.
- To evaluate the effectiveness of Cword2vec in Urdu sentiment analysis using deep learning models.
Main Methods:
- Developed a self-trained morphological rule-based compound word embedding (Cword2vec) model based on word2vec.
- Applied Cword2vec for text representation in Urdu sentiment analysis tasks.
- Evaluated Cword2vec performance using Long Short-Term Memory (LSTM), Bidirectional LSTM (BiLSTM), Convolutional Neural Networks (CNN), and Convolutional LSTM (C-LSTM).
- Compared Cword2vec against traditional bigram and trigram approaches for compound word identification.
Main Results:
- The proposed Cword2vec model demonstrated superior performance across all evaluated deep learning models.
- Morphological rule-based compound word embeddings significantly outperformed traditional bigram and trigram methods.
- Improvements were observed in key metrics including precision, recall, F1 score, and accuracy.
Conclusions:
- Morphological rule-based compound word embeddings offer a more effective approach for Urdu NLP tasks compared to traditional methods.
- Cword2vec provides a robust solution for handling Urdu's complex word structures, enhancing sentiment analysis.
- Further research into Urdu NLP, particularly with morphological insights, is warranted to leverage its linguistic richness.
Related Concept Videos
UV–Vis Spectroscopy: Woodward–Fieser Rules
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Degree of Unsaturation
The degree of unsaturation for hydrocarbons is U = (2C + 2 − H) / 2, where C is the number of carbon atoms and H is the number of hydrogen atoms.
Empathy
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
