基准测试Hook and Bait乌尔都语新闻数据集用于使用大型语言模型检测域位和多语言假新闻
Sheetal Harris1, Jinshuo Liu2, Hassan Jalil Hadi3
1School of Cyber Science and Engineering, Wuhan University, Wuhan, China.
Scientific reports
|May 3, 2025
概括
这项研究引入了一个新的乌尔都语数据集,并使用大型语言模型 (LLM) 来有效地检测乌尔都语和英语的假新闻 (FND). 该方法实现了高精度,为资源较少的语言提供了解决方案.
科学领域:
- 自然语言处理自然语言处理.
- 人工智能的人工智能
- 计算语言学 计算语言学
背景情况:
- 检测假新闻 (FN) 是一个全球性的挑战,特别是在资源较低的语言中,因为注释数据有限.
- 现有的假新闻检测 (FND) 多语言方法对于资源稀缺的语言是不够的.
- 大型语言模型 (LLM) 为推进多语言FND提供了一个有希望的途径.
研究的目的:
- 开发和评估使用大型语言模型 (LLM) 进行假新闻检测 (FND) 的自动化机制.
- 为FND研究策划第一个大型的,多领域的乌尔都语新闻库.
- 评估基于LLM的FND在单语言 (乌尔都语) 和多语言 (乌尔都语和英语) 环境中的表现.
主要方法:
- 策划了"Hook and Bait Urdu"集体,包括78,409篇真假新闻文章.
- 在乌尔都语的体上微调了LLaMA 2模型,用于单模FND.
- 采用基于LLaMA 2的框架,精心调整了乌尔都语和英语数据集,用于多语言FND.
- 使用了LoRA微调方法,优化了超参数和早期停止以提高效率.
主要成果:
- 实现了0.978的准确性和0.971的F1得分为单模乌尔都语FND.
- 证明了强大的多语言FND性能,准确度为0.984和F1得分为0.980.
- 轻量级的LoRA方法确保了以最小可训练参数 (0.032%) 的计算效率.
结论:
- 提出的基于LLM的方法显著提升了自动虚假新闻检测 (FND),特别是在低资源语言.
- "Hook and Bait Urdu"数据集和精心调整的LLaMA 2模型为单语言和多语言FND提供了一个强大的框架.
- 公开可用的数据集有助于进一步研究和开发FND机制来打击错误信息.
相关概念视频
Difference from Background: Limit of Detection
4.6K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
4.6K
Detection of Black Holes
2.1K
Although black holes were theoretically postulated in the 1920s, they remained outside the domain of observational astronomy until the 1970s.
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
2.1K
Improving Translational Accuracy
2.5K
2.5K
Bias
3.7K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.7K
Detection of Gross Error: The Q Test
4.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
4.4K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K


