使用ADHAR检测仇恨言论:阿拉伯语中的多方言仇恨言论集体
Anis Charfi1, Mabrouka Besghaier1, Raghda Akasheh1
1Information Systems Department, Carnegie Mellon University, Doha, Qatar.
Frontiers in artificial intelligence
|June 14, 2024
概括
本研究介绍了ADHAR,这是一个平衡的阿拉伯仇恨言论数据集,涵盖多种方言和类别. 它使强大的仇恨言论检测模型能够实现高达95%的准确性.
科学领域:
- 自然语言处理自然语言处理.
- 计算语言学 计算语言学
- 人工智能的人工智能
背景情况:
- 由于方言的多样性,阿拉伯语仇恨言论的检测具有挑战性.
- 现有的数据集范围有限,缺乏方言和分类平衡.
- 这就需要为有效的模型培训提供全面和平衡的资源.
研究的目的:
- 介绍ADHAR,一个新的,全面的,多方言,多类别的阿拉伯语仇恨言论集体.
- 通过确保方言,类别和仇恨/非仇恨类别之间的平衡来解决现有数据集的局限性.
- 促进对阿拉伯仇恨言论检测模型的公正评估和开发.
主要方法:
- 在现代标准阿拉伯语 (MSA) 和五种主要方言 (埃及语,黎凡提语,海湾语,马格里比语) 中系统地收集数据.
- 严格的注释过程,涉及每个方言的多个注释者,用于四个仇恨言论类别 (国籍,宗教,种族,种族).
- 精心策划的数据集 (70,369个单词) 以确保跨方言,类别和情绪的平衡.
主要成果:
- 通过广泛的分析,ADHAR集体证明了其高质量和实用性.
- 实验表明,经典和深度学习模型在仇恨言论和类别检测方面达到高达90%的准确性和92%的F1分数.
- 与Arabert进行的培训分别实现了94%的准确性和95%的F1-分数,分别用于仇恨言论和类别检测.
结论:
- 阿达尔是一个有价值的,平衡的资源,用于推进阿拉伯的仇恨言论检测.
- 该数据集使得开发强大而准确的仇恨言论分类器成为可能.
- 未来的研究可以利用ADHAR进行改进的跨方言和多类别的仇恨言论分析.
相关概念视频
Detection of Black Holes
2.2K
Although black holes were theoretically postulated in the 1920s, they remained outside the domain of observational astronomy until the 1970s.
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
2.2K
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Bias
4.2K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
4.2K
Classification of Signals
437
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
437
Stereotypes, Prejudice, and Discrimination
90.1K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
90.1K
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K


