相关实验视频
Updated: Jan 13, 2026

13:19
Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
9.9K
ViClickbait-2025:越南人点击诱检测的综合数据集
Dai Phuoc Nguyen1,2, Thien Khai Tran3, Y Minh Nguyen2
1Faculty of Information Technology, HUTECH University, Vietnam.
Data in brief
|October 29, 2025
概括
研究人员开发了ViClickbait-2025,这是一个用于自动检测点击诱的越南数据集. 该资源有助于识别欺骗性的在线标题,提高内容的可信度.
科学领域:
- 自然语言处理自然语言处理.
- 机器学习 机器学习
- 信息检索 信息检索
背景情况:
- 点击诱惑标题在在线信息消费方面带来了挑战.
- 开发强大的自动点击诱检测系统对于内容可信度至关重要.
- 需要在非英语语言 (如越南语) 中专门的数据集.
研究的目的:
- 介绍ViClickbait-2025,一个新的越南语数据集用于点击诱检测研究.
- 为培训和评估用于点击诱识别的机器学习模型提供全面的资源.
- 促进理解和打击欺骗性的在线内容的进步.
主要方法:
- 来自八个越南新闻平台的3414个头条新闻 (2023-2025年) 的网络扫描.
- 标题作为点击诱惑或非点击诱惑的注释由三个独立评论员 (科恩·卡帕:0.822).
- 数据预处理包括HTML删除,除重复,规范化和包含九个关键属性.
主要成果:
- "ViClickbait-2025"数据集包含了13个新闻类别中的31.2%的点击诱标题.
- 高度的注释者间协议 (科恩的卡帕=0.822) 确保了注释质量.
- 数据集在CC BY 4.0许可证下以JSONL和CSV格式提供.
结论:
- ViClickbait-2025是越南点击诱检测研究的一个有价值的,高质量的资源.
- 该数据集支持开发更准确,更可靠的自动点击诱检测模型.
- 这项工作有助于提高越南在线新闻消费的可信度.
相关概念视频
Difference from Background: Limit of Detection
8.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
8.0K
Detection of Gross Error: The Q Test
6.9K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.9K
Detection of Black Holes
2.5K
Although black holes were theoretically postulated in the 1920s, they remained outside the domain of observational astronomy until the 1970s.
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
2.5K
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K
Relative Frequency Histogram
6.3K
The relative frequency depicts the proportion of data points that have each value. The frequency tells the number of data points that have each value. Like the histogram, a relative frequency histogram also has the same shape with a horizontal scale (the x-axis), but the vertical scale (the y-axis) is marked with relative frequencies (percentages of the whole) instead of actual frequencies. A relative frequency histogram is a graphical representation of a frequency distribution where the...
6.3K
Probability Histograms
13.1K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
13.1K