Related Experiment Video
Updated: Jun 30, 2025

06:48
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.2K
A normalization model for repeated letters in social media hate speech text based on rules and spelling correction
Zainab Mansur1, Nazlia Omar1, Sabrina Tiun1
1Center for AI Technology (CAIT), FTSM, Universiti Kebangsaan Malaysia, UKM, Bangi, Malaysia.
Plos One
|March 21, 2024
Summary
This study introduces an unsupervised model to normalize words with repeated letters, reducing out-of-vocabulary (OOV) instances and improving hate speech detection accuracy by correctly replacing problematic terms.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Social Media Analysis
Background:
- Online hate speech is increasing with social media use.
- Words with repeated letters contribute to out-of-vocabulary (OOV) issues in hate speech detection.
- Existing models struggle to correctly normalize OOV words with repeated letters.
Purpose of the Study:
- To develop an improved unsupervised model for normalizing OOV words with repeated letters.
- To enhance the accuracy of hate speech detection by improving text normalization.
- To replace OOV words with correct in-vocabulary (IV) alternatives.
Main Methods:
- Combined rule-based patterns for repeated letters with the SymSpell algorithm.
- Developed rules based on letter repetition position (beginning, middle, end) and pattern.
- Utilized an unsupervised approach, avoiding special dictionaries or annotated data.
Main Results:
- Reduced the percentage of OOV words to 8%.
- Achieved F1 scores 9% and 13% higher than two benchmark studies.
- Demonstrated superior performance in replacing OOV words with correct IV replacements.
Conclusions:
- Rule-based patterns combined with spelling correction effectively normalize words with repeated letters.
- The proposed normalization model significantly improves hate speech detection performance.
- This unsupervised method offers a viable solution for enhancing text normalization in NLP tasks.
Related Concept Videos
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K
Language
213
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
213
Components of Language
269
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
269
Learning Disabilities
112
Learning disabilities are cognitive disorders caused by neurological impairments that affect cognitive functions like language and reading, without indicating overall intellectual or developmental challenges. These disabilities differ from global intellectual or developmental disabilities as they are limited to distinct cognitive functions. Common learning disabilities include dysgraphia, dyslexia, and dyscalculia, each of which impacts unique aspects of learning.
Dyslexia
Dyslexia is a...
Dyslexia
Dyslexia is a...
112
Mismatch Repair
40.1K
Overview
40.1K
Sign Test for Matched Pairs
131
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in...
To conduct the sign test, we first calculate the differences in...
131

