Related Experiment Video
Updated: Sep 25, 2025

06:37
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
4.2K
Detecting racism and xenophobia using deep learning models on Twitter data: CNN, LSTM and BERT
José Alberto Benítez-Andrades1, Álvaro González-Jiménez2, Álvaro López-Brea2
1SALBIS Research Group, Department of Electric, Systems and Automatics Engineering, Universidad de León, León, León, Spain.
Peerj. Computer Science
|May 2, 2022
Summary
A new BERT-based model, BETO, excels at detecting racist and xenophobic tweets in Spanish. This natural language processing advancement highlights the need for Spanish-specific training data in AI content moderation.
Area of Science:
- Computational Linguistics
- Artificial Intelligence
- Social Media Analysis
Background:
- Manual content moderation on social media is infeasible due to rapid growth.
- Natural Language Processing (NLP) enables automated text classification.
- Existing NLP models often lack sufficient training data for specific languages like Spanish.
Purpose of the Study:
- To develop and compare BERT-based and other deep learning models for detecting racist and xenophobic messages in Spanish tweets.
- To evaluate the performance of models trained on Spanish-specific data versus multilingual data.
- To underscore the importance of native transfer learning models for Spanish NLP tasks.
Main Methods:
- Development of five predictive models: two BERT-based (BETO, mBERT) and three deep learning models (CNN, LSTM, CNN+LSTM).
- Training and evaluation of models using Spanish tweets containing racist and xenophobic content.
- Comparative analysis of model performance based on precision metrics.
Main Results:
- The BETO model, a BERT variant trained exclusively on Spanish text, achieved the highest precision at 85.22%.
- The mBERT model achieved 82.00% precision, outperforming other deep learning models.
- CNN, LSTM, and CNN+LSTM models showed precision ranging from 79.34% to 80.48%.
Conclusions:
- Native transfer learning models, like BETO, are crucial for effectively addressing NLP challenges in Spanish.
- The study demonstrates the superior performance of Spanish-specific models in detecting hate speech.
- Findings support the development of tailored NLP solutions for diverse linguistic contexts.
Related Concept Videos
Stereotypes, Prejudice, and Discrimination
91.8K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
91.8K
Stereotype Content Model
14.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.9K

