Related Experiment Video
Updated: Jan 7, 2026

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Bengali cyberbullying detection: A comprehensive dataset for advanced analysis
Gazi Tahsina Sharmin Jahin1, Firoze Maliha1, Nuren Nafisa1
1Department of Computer Science and Engineering, International Islamic University Chittagong, Kumira, IIUC Avenue, Sonaichhari, Chattogram, 4318, Chattogram, Bangladesh.
This study analyzes Bengali cyberbullying on social media, identifying key themes and achieving 93% accuracy with advanced models like mBERT and hybrid CNN approaches. Local Interpretable Model-agnostic Explanations enhances model transparency.
Area of Science:
- Computational Linguistics
- Social Computing
- Cyberpsychology
Background:
- Cyberbullying is a significant digital concern, amplified by social media.
- Research on cyberbullying is extensive, yet studies on Bengali cyberbullying are limited.
- This paper addresses the scarcity of research on Bengali cyberbullying.
Purpose of the Study:
- To analyze Bengali cyberbullying on social media platforms.
- To identify prevalent themes and features of negative Bengali comments.
- To evaluate and compare the performance of various machine learning models for cyberbullying detection.
Main Methods:
- Analysis of over 70,000 Bengali social media comments.
- Sentiment analysis to classify comments as positive or negative.
- Latent Dirichlet Allocation (LDA) for topic modeling of negative comments.
- Application and comparison of machine learning models: Support Vector Machine, XGBoost, CNN+BiLSTM+GRU, mBERT, XLM-R.
- Integration of BERT embeddings with CNN and Artificial Neural Network (ANN) models.
- Utilizing Local Interpretable Model-agnostic Explanations (LIME) for model interpretability.
Main Results:
- mBERT achieved the highest accuracy at 92%.
- The CNN+BiLSTM+GRU hybrid model reached 91% accuracy.
- Incorporating BERT embeddings into CNN and ANN models improved performance to 93% accuracy.
- LDA successfully extracted features related to age, gender, ethnicity, religion, and miscellaneous categories from negative comments.
Conclusions:
- Advanced deep learning models, particularly mBERT and hybrid architectures, are effective for detecting Bengali cyberbullying.
- BERT embeddings significantly enhance the performance of cyberbullying detection models.
- LIME provides valuable insights into model predictions, increasing transparency and trust in cyberbullying detection systems.
- This research contributes a robust framework for understanding and combating Bengali cyberbullying.
Related Concept Videos
Bullying
Mass Analyzers: Overview
Mass Analyzers: Common Types