Related Experiment Video
Updated: Aug 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Improving the Polarity of Text through word2vec Embedding for Primary Classical Arabic Sentiment Analysis
Nour Elhouda Aoumeur1,2, Zhiyong Li1,2, Eissa M Alshari3
1College of Computer Science and Electronic Engineering, Hunan University, Lushan, Changsha, 410082 Hunan China.
This study introduces a new dataset for sentiment analysis of classical Arabic literature, utilizing advanced word embedding techniques. Logistic Regression with Word2Vec achieved the highest accuracy in predicting topic-polarity.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Machine Learning
Background:
- Sentiment analysis research has grown significantly, but classical Arabic literature remains under-explored.
- Existing Arabic text feature extraction methods often rely on word frequency, neglecting semantic relationships.
- No dedicated dataset or advanced feature extraction for sentiment analysis of classical Arabic books was available.
Purpose of the Study:
- To develop a novel classical Arabic dataset (CASAD) for sentiment analysis from art books.
- To implement advanced word embedding techniques for extracting deep semantic features from formal Arabic text.
- To evaluate the effectiveness of various machine learning models for sentiment classification on this dataset.
Main Methods:
- Creation of the CASAD dataset by collecting and human-expert labeling sentences from classical Arabic art books.
- Application of word embedding techniques (similar to Word2Vec) to extract semantic features.
- Evaluation of features using machine learning algorithms: Support Vector Machines (SVM), Logistic Regression (LR), Naive Bayes (NB), K-Nearest Neighbors (KNN), Latent Dirichlet Allocation (LDA), and Classification And Regression Trees (CART).
- Statistical validation and reliability testing of dataset labels.
- Tenfold cross-validation to assess classification rates for two and three classes.
Main Results:
- The Logistic Regression model combined with Word2Vec embeddings demonstrated superior performance in sentiment classification.
- The developed CASAD dataset and feature extraction methods provide a robust foundation for Arabic sentiment analysis.
- Comparative analysis showed varying classification accuracies across different machine learning algorithms.
Conclusions:
- The Logistic Regression with Word2Vec approach is highly effective for sentiment analysis in classical Arabic.
- The CASAD dataset and proposed feature extraction method advance the field of Arabic Natural Language Processing.
- This research provides a benchmark for future sentiment analysis studies on classical Arabic texts.
Related Concept Videos
Group Polarization
Bond Polarity, Dipole Moment, and Percent Ionic Character
Electrophilic Aromatic Substitution: Sulfonation of Benzene
Valence Bond Theory
Molecular Shape and Polarity
Polymer Classification: Stereospecificity

