为了更好的希伯来语点击诱检测:从BERT和数据增强的洞察力
Talya Natanya1, Chaya Liebeskind1
1Department of Computer Science, Jerusalem College of Technology, Jerusalem, Israel.
PloS one
|November 6, 2025
概括
这项研究使用深度学习和数据增强来增强希伯来语点击诱检测,达到92%的准确性. 这些方法提高了内容质量和用户对数字媒体的信任.
科学领域:
- 自然语言处理自然语言处理.
- 机器学习 机器学习
- 数字媒体分析 数字媒体分析
背景情况:
- 点击诱惑标题传播错误信息,破坏在线内容的可信度.
- 准确的点击诱检测对于信息质量和用户信任至关重要.
- 希伯来语是一个资源较低的语言,对点击诱检测提出了独特的挑战.
研究的目的:
- 通过深度学习推进希伯来语点击诱检测.
- 探索各种数据增强策略的影响.
- 在点击诱识别中超越以前的准确性基准.
主要方法:
- 使用基于BERT的深度学习模型.
- 实施各种数据增强技术 (弱监督,替代,生成,基于语言).
- 应用了对最先进的希伯来语言模型的增强.
主要成果:
- 有针对性的增强,特别是词级和上下文增强,提高了性能.
- 实现了92%的最高准确度,超过了传统的机器学习方法 (87%).
- 证明了将深度学习与定制数据增强相结合的有效性.
结论:
- 深度学习和数据增强显著提高了希伯来语点击诱检测.
- 开发的框架可以应用于现实世界的系统,以提高内容质量.
- 提供了一个可复制的模型,用于在其他代表性不足的语言中检测点击诱.
相关概念视频
Detection of Black Holes
2.5K
Although black holes were theoretically postulated in the 1920s, they remained outside the domain of observational astronomy until the 1970s.
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
2.5K
Detection of Gross Error: The Q Test
6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K
Difference from Background: Limit of Detection
8.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
8.0K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
