维拉-阿拉布:通过构建平衡的新闻数据集来进行真实性分析,揭示阿拉伯语推文的可信性
Mohamed A Mostafa1, Ahmad Almogren2
1Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia.
PeerJ. Computer science
|December 9, 2024
概括
VERA-ARAB是一个新的数据集,用于检测阿拉伯语推特中的假新闻. 本资源有助于改进机器学习模型,用于社交媒体上的真实性分析.
科学领域:
- 自然语言处理自然语言处理.
- 社交媒体分析 社交媒体分析
- 计算语言学 计算语言学
背景情况:
- 社交媒体上错误信息的增加需要强大的工具来检测假新闻.
- 现有的数据集往往缺乏有效的阿拉伯假新闻分析所需的规模,多样性或特定的语言重点.
更多相关视频
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
502
07:12Protocol for Data Collection and Analysis Applied to Automated Facial Expression Analysis Technology and Temporal Analysis for Sensory Evaluation
Published on: August 26, 2016
9.4K
相关概念视频
Accuracy and Precision
8.7K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value. Highly accurate...
8.7K
Improving Translational Accuracy
2.5K
2.5K
Margin of Error
4.0K
The margin of error is also called the maximum error of an estimate. The margin of error is the maximum possible or expected difference between the observed sample parameter value and the actual population parameter value. For proportion, it is the maximum difference between the value of sample proportion obtained from the data and the true value of population proportion. As the true value of the population parameter is not known, the margin of error is calculated using the sample statistic.
4.0K
Root Mean Square
3.2K
If in an experiment, data values have a probability of being both positive and negative, neither the arithmetic mean, the geometric mean, nor the harmonic mean can be used to calculate the central tendency of the data set. In particular, if the positive and negative values are equally likely, the arithmetic mean is close to zero.
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
3.2K
Two-Way ANOVA
2.6K
The two-way ANOVA is an extension of the one-way ANOVA. It is a statistical test performed on three or more samples categorized by two factors - a row factor and a column factor. Ronald Fischer mentioned it in 1925 in his book 'Statistical Methods for Researchers.'
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
2.6K
Arithmetic Mean
13.4K
The arithmetic mean is the most commonly used measure of the central tendency of a data set. It is defined as the sum of all the elements constituting the data set, divided by the total number of elements. It is sometimes loosely referred to as the “average.”
When all the values in a data set are not unique, the sum in the numerator can be calculated by multiplying each distinct value by its frequency.
Sometimes, the arithmetic mean of a sample can be affected by a few data points...
When all the values in a data set are not unique, the sum in the numerator can be calculated by multiplying each distinct value by its frequency.
Sometimes, the arithmetic mean of a sample can be affected by a few data points...
13.4K
