Related Experiment Video
Updated: May 20, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
474
Data augmentation for Arabic text classification: a review of current methods, challenges and prospective directions
Samia F Abdhood1,2, Nazlia Omar1, Sabrina Tiun1
1Center for Artificial Intelligence Technology, Faculty of Information Science and Technology, Universiti Kebangsaan Malaysia, Bangi, Selangor, Malaysia.
Peerj. Computer Science
|March 26, 2025
Summary
Data augmentation enhances Arabic text classification by artificially creating data to overcome dataset limitations. Research primarily focuses on sentiment analysis and propaganda detection, with a need for more studies on diverse text types and datasets.
Area of Science:
- Natural Language Processing (NLP)
- Machine Learning
- Artificial Intelligence
Background:
- Data augmentation techniques artificially generate new data to improve machine learning model performance.
- These methods address challenges like limited training data and class imbalance in various domains.
- Their application in Arabic text classification is crucial for enhancing classifier performance.
Purpose of the Study:
- To review and analyze data augmentation techniques specifically for Arabic text classification.
- To provide a comprehensive understanding of existing approaches in Arabic Natural Language Processing (ANLP).
- To identify research trends, gaps, and future directions in this specialized field.
Main Methods:
- A systematic review of Arabic studies on data augmentation in text classification published between 2019 and 2024.
- Application of specific inclusion and exclusion criteria to ensure a focused and relevant analysis.
- Synthesis of findings regarding dominant research areas, text types, and dataset availability.
Main Results:
- Research predominantly covers sentiment analysis and propaganda detection, with limited exploration in areas like sarcasm detection.
- A significant lack of benchmark datasets for Arabic text classification tasks was observed.
- Most studies concentrate on short texts (e.g., social media), highlighting a need for research on long texts.
Conclusions:
- Further investigation into data augmentation for Arabic long texts is essential.
- A rigorous comparison of techniques tailored to Arabic's unique linguistic features is required.
- This review offers valuable insights for advancing Arabic NLP and selecting optimal data augmentation strategies.

