在低资源语言中检测攻击性语言:波斯语语言的用例
Marzieh Mozafari1, Khouloud Mnassri1, Reza Farahbakhsh1
1Samovar, Télécom SudParis, Institut Polytechnique de Paris, Palaiseau, France.
PloS one
|June 21, 2024
概括
这项研究引入了新的波斯攻击性语言数据集和用于检测社交媒体上滥用内容的模型. 综合模型显著提高了针对个人或群体的冒犯性语言的检测准确度.
科学领域:
- 计算语言学 计算语言学
- 自然语言处理自然语言处理.
- 社交媒体分析 社交媒体分析
背景情况:
- 滥用内容,包括仇恨言论和冒犯性语言,在社交媒体平台上普遍存在.
- 检测冒犯性语言是具有挑战性的,特别是在低资源语言中,因为注释数据有限.
- 现有的研究主要集中在像英语这样的高资源语言上,忽视了其他语言.
研究的目的:
- 为了解决缺乏资源的攻击性语言检测在波斯语,一个低资源的语言.
- 从社交媒体创建和注释一本新的波斯攻击性语言库.
- 开发和评估用于检测波斯攻击性语言的机器学习和深度学习模型.
主要方法:
- 创建了一个由6,000个波斯微博帖子组成的新群体,并对攻击性语言及其目标进行注释.
- 使用经典机器学习 (ML),深度学习 (DL) 和基于变压器的模型 (例如,ParsBERT) 进行了实验.
- 通过整合多个模型来提高检测性能,提出了一个合体模型.
主要成果:
- 使用n-gram和ParsBERT模型的支持向量机 (SVM) 在初始单模型评估中表现强.
- 拟议的堆叠组合模型显著优于单个模型的性能.
- 整体模型在三个注释级别 (进攻性与非进攻性,有针对性与非有针对性,个人与群体) 中实现了5%的宏观F1得分改善.
结论:
- 开发的波斯攻击性语言群体是低资源语言NLP研究的宝贵资源.
- 组合方法可以有效地提高攻击性语言检测系统的性能.
- 这些发现有助于减轻波斯语社交媒体社区在线伤害.
相关概念视频
Types of Errors: Detection and Minimization
1.6K
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
1.6K
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Censoring Survival Data
78
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
78
Nonsense-mediated mRNA Decay
10.6K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
10.6K
Extraction: Advanced Methods
446
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
446
Difference from Background: Limit of Detection
6.3K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.3K


