跨语言的抄袭检测:对阿拉伯语和英语到阿拉伯语的长文件进行全面研究
Ahmad Abdelaal1, Abdallah Elsaadany1, Abdelrhman Ahmed Medhat1
1School of Information Technology and Computer Science, Nile University, Giza, Egypt.
PeerJ. Computer science
|September 24, 2025
概括
本研究介绍了一种先进的阿拉伯抄袭检测系统,使用罗神经网络 (SNN) 和像AraT5.5这样的变压器模型. 这种新的方法显著提高了准确性,在检测文本相似性方面达到0.9058 F1分.
科学领域:
- 自然语言处理自然语言处理.
- 计算语言学 计算语言学
- 人工智能的人工智能
背景情况:
- 由于语言复杂性和数据稀缺性,阿拉伯语抄袭检测面临着挑战.
- 现有的方法在阿拉伯文文本中与细微的语义和结构变化作斗争.
研究的目的:
- 开发一个强大的框架,用于阿拉伯抄袭的检测.
- 为了提高识别阿拉伯语中类似文本的准确性.
主要方法:
- 整合罗神经网络 (SNN) 与变压器架构 (AraT5,长变压器).
- 混合工作流 结合变压器编码器和分类目标.
- 利用加权交叉损失和子损失来处理不平衡的数据集.
主要成果:
- 拟议的具有加权交叉损失的AraT5实现了0.9058.5的高F1得分.
- 与现有方法相比,该系统表现出优越的性能.
- 在ExAraCorpusPAN2015数据集上验证了有效性.
结论:
- 基于变压器的架构和特定类的损失函数对于改善阿拉伯语抄袭检测至关重要.
- 拟议的框架为资源不足的语言提供了显著的进步.
- 突出了深度学习在阿拉伯语中进行复杂文本分析的潜力.
相关概念视频
Proofreading
8.7K
Synthesis of new DNA molecules is carried out by the enzyme DNA polymerase, which adds nucleotides on the daughter strand complementary to the template DNA strand. DNA polymerase has a higher affinity to add the correct base and ensures fidelity during DNA replication. Furthermore, it exhibits proofreading activity during replication, using an exonuclease domain that cuts off incorrect nucleotides from the nascent DNA strand.
Errors During Replication are Corrected by the DNA Polymerase...
Errors During Replication are Corrected by the DNA Polymerase...
8.7K
Proofreading
60.0K
Overview
60.0K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
Detection of Gross Error: The Q Test
6.9K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.9K
Types of Errors: Detection and Minimization
10.0K
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
10.0K


