Related Experiment Video
Updated: Aug 7, 2025

Continuous Theta Burst Stimulation of the Posterior Medial Frontal Cortex to Experimentally Reduce Ideological Threat Responses
Published on: September 28, 2018
Multi-Ideology, Multiclass Online Extremism Dataset, and Its Evaluation Using Machine Learning
Mayur Gaikwad1, Swati Ahirrao1, Shraddha Phansalkar2
1Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune, MH 412115, India.
This study introduces a new, balanced dataset for detecting extremism online, covering multiple ideologies and classifying text into propaganda, radicalization, and recruitment. Machine learning models, particularly support vector machines, show promising results for this multi-class extremism detection task.
Area of Science:
- Computer Science
- Social Sciences
- Computational Linguistics
Background:
- Social media platforms are exploited for extremist propaganda, radicalization, and recruitment, necessitating effective detection methods.
- Existing extremism detection research is limited by single-ideology datasets, binary classification, and manual data validation.
- Challenges include class imbalance and a lack of automated data validation in current extremism detection studies.
Purpose of the Study:
- To develop a balanced, multi-ideology extremism text dataset for improved detection.
- To classify extremism into propaganda, radicalization, and recruitment categories using robust validation.
- To address limitations of existing single-ideology and binary classification datasets.
Main Methods:
- Created a versatile extremism text dataset generalizing multiple ideologies (ISIS, White Supremacist).
- Utilized TF-IDF (unigram, bigrams, trigrams) and pretrained word2vec features for analysis.
- Evaluated machine learning classifiers including Naïve Bayes, Support Vector Machine, Random Forest, and XGBoost.
Main Results:
- The Support Vector Machine with TF-IDF unigram features achieved the best performance with a 0.67 F1 score.
- The proposed multi-ideology, multi-class dataset demonstrated comparable performance to existing single-ideology datasets.
- Feature extraction included TF-IDF and word2vec for semantic analysis.
Conclusions:
- The developed multi-ideology dataset enhances extremism text classification accuracy.
- Machine learning models can effectively identify propaganda, radicalization, and recruitment on social media.
- This research contributes a valuable resource for combating online extremism through improved detection.
Related Concept Videos
Group Polarization
Stereotype Content Model
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Multi-input and Multi-variable systems
In the absence...

