Related Experiment Video
Updated: Jul 8, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Pashto offensive language detection: a benchmark dataset and monolingual Pashto BERT
Ijazul Haq1, Weidong Qiu1, Jie Guo1
1School of Cyber Science and Engineering, Shanghai Jiao Tong University, Shanghai, Minhang, China.
This study introduces the Pashto Offensive Language Dataset (POLD) for detecting offensive content in Pashto social media. A Pashto BERT model achieved high accuracy, outperforming other AI methods for this low-resource language.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Artificial Intelligence
Background:
- Offensive language on social media poses a threat to online communities.
- Detecting offensive content in low-resource languages like Pashto is an underexplored research area.
- Existing research primarily focuses on high-resource languages, leaving a gap for Pashto.
Purpose of the Study:
- To develop an AI model for automatic offensive text detection in the Pashto language.
- To create a benchmark dataset for Pashto offensive language identification.
- To evaluate and compare various deep learning and transfer learning models for this task.
Main Methods:
- Developed the Pashto Offensive Language Dataset (POLD) from Twitter data.
- Implemented and evaluated deep learning models (CNNs, RNNs) with static word embeddings (Word2Vec, fastText, GloVe).
- Investigated transfer learning using XLM-R and a custom-trained Pashto BERT model.
Main Results:
- The Pashto BERT model achieved superior performance compared to other evaluated models.
- The Pashto BERT model attained an F1-score of 94.34% and an accuracy of 94.77%.
- Deep learning and transfer learning approaches were benchmarked on the POLD dataset.
Conclusions:
- The developed Pashto BERT model is effective for offensive text detection in Pashto.
- This research contributes a valuable dataset and a high-performing model for a low-resource language.
- Automatic detection of offensive content in Pashto is feasible and crucial for online safety.
Related Concept Videos
Detection of Gross Error: The Q Test
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Quantifying and Rejecting Outliers: The Grubbs Test
Detection of Black Holes
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Margin of Error

