基于规则和拼写纠正的社交媒体仇恨言论文本中重复出现的字母的正常化模型
Zainab Mansur1, Nazlia Omar1, Sabrina Tiun1
1Center for AI Technology (CAIT), FTSM, Universiti Kebangsaan Malaysia, UKM, Bangi, Malaysia.
PloS one
|March 21, 2024
概括
这项研究引入了一个无监督的模型,以使重复字母的单词正常化,减少词汇库外的 (OOV) 实例,并通过正确替换有问题的术语来提高仇恨言论检测准确度.
科学领域:
- 自然语言处理自然语言处理.
- 计算语言学 计算语言学
- 社交媒体分析 社交媒体分析
背景情况:
- 在线仇恨言论随着社交媒体的使用而增加.
- 具有重复字母的单词有助于在仇恨言论检测中发现词汇库外 (OOV) 问题.
- 现有的模型很难正确地将重复字母的OOV单词正常化.
研究的目的:
- 开发一个改进的无监督模型来规范重复字母的OOV单词.
- 通过改进文本规范化来提高仇恨言论检测的准确性.
- 将OOV单词替换为正确的词汇 (IV) 替代品.
主要方法:
- 结合基于规则的模式重复的字母与SymSpell算法.
- 根据字母重复位置 (开始,中间,结束) 和模式制定规则.
- 使用无监督的方法,避免使用特殊的字典或注释数据.
主要成果:
- 将OOV单词的百分比降低到8%.
- 获得的F1分数比两个基准研究分别高出9%和13%.
- 在用正确的IV替换取代OOV单词方面表现出卓越的性能.
结论:
- 基于规则的模式与拼写纠正相结合,有效地使重复字母的单词正常化.
- 拟议的规范化模型显著提高了仇恨言论检测性能.
- 这种无监督的方法为增强NLP任务中的文本规范化提供了可行的解决方案.
相关概念视频
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K
Language
213
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
213
Components of Language
269
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
269
Learning Disabilities
112
Learning disabilities are cognitive disorders caused by neurological impairments that affect cognitive functions like language and reading, without indicating overall intellectual or developmental challenges. These disabilities differ from global intellectual or developmental disabilities as they are limited to distinct cognitive functions. Common learning disabilities include dysgraphia, dyslexia, and dyscalculia, each of which impacts unique aspects of learning.
Dyslexia
Dyslexia is a...
Dyslexia
Dyslexia is a...
112
Mismatch Repair
40.1K
Overview
40.1K
Sign Test for Matched Pairs
131
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in...
To conduct the sign test, we first calculate the differences in...
131


