Related Experiment Video
Updated: Jan 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Leveraging pre-trained embeddings in an ensemble machine learning approach for Arabic sentiment analysis.
Areej Jaber1, Israa Bahati1, Paloma Martínez2
1Computer Science Department, Palestine Technical University - Kadoorie, Tulkarm, Palestine.
Frontiers in Artificial Intelligence
|September 29, 2025
Summary
Ensemble machine learning methods significantly improve Arabic sentiment analysis, outperforming individual classifiers. These approaches effectively handle linguistic complexities and imbalanced datasets, enhancing system robustness and generalizability.
Area of Science:
- Natural Language Processing
- Machine Learning
- Computational Linguistics
Background:
- Arabic sentiment analysis faces challenges due to linguistic diversity, dialectal variations, and limited resources.
- Developing robust sentiment classification systems requires addressing these inherent complexities.
Purpose of the Study:
- To investigate the effectiveness of ensemble machine learning methods for Arabic sentiment analysis.
- To evaluate homogeneous ensemble techniques on both balanced and imbalanced Arabic datasets.
- To assess the impact of pre-trained word embeddings and SMOTE on model performance.
Main Methods:
- Implementation and evaluation of homogeneous ensemble techniques (e.g., Naive Bayes, SVM, Decision Tree, SGD, KNN, Random Forest).
- Utilized two datasets: ArTwitter (balanced) and Syria_Tweets (imbalanced).
- Employed Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalance and incorporated pre-trained word embeddings and unigram features.
Main Results:
- Ensemble models consistently outperformed individual classifiers across both datasets.
- On ArTwitter, an ensemble achieved 90.22% accuracy and 92.0% F1-score.
- On Syria_Tweets, another ensemble reached 83.82% accuracy and 83.86% F1-score.
Conclusions:
- Ensemble learning enhances the robustness and generalizability of Arabic sentiment analysis systems.
- Pre-trained embeddings further boost performance, demonstrating the value of these approaches.
- Ensemble methods effectively overcome challenges in Arabic NLP, including linguistic complexity and data imbalance.
Related Concept Videos
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
Surveys
16.6K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
16.6K
Aggregates Classification
970
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
970
