同时发生的新闻媒体热点挖矿-文本挖矿方法设计的单词模型
1School of Arts and Creative Technologies, The University of York, York, United Kingdom.
Mathematical biosciences and engineering : MBE
|June 14, 2024
概括
这项研究引入了新闻媒体热点挖掘的新型共发生词模型,显著提高了识别热门话题的准确性和效率. 该方法增强了数字时代的及时信息发现.
科学领域:
- 自然语言处理自然语言处理.
- 数据挖掘 数据挖掘
- 计算语言学 计算语言学
背景情况:
- 在线媒体的扩散需要先进的方法来识别热门话题.
- 传统的热点挖掘算法缺乏当代新闻分析所需的精度和速度.
- 现有的方法在捕捉新兴新闻趋势时,往往在准确性和及时性方面扎.
研究的目的:
- 为了解决当前新闻媒体热点挖掘技术的局限性.
- 提出一个改进的热点采矿方法,以提高准确性和及时性.
- 为高效的数据处理利用计算框架.
主要方法:
- 开发一种新的共同出现词模型,其中包含单词权重.
- 使用共发生模型和改进的平滑反向频率排名 (SIFRANK) 的热点采矿算法的设计.
- 集成Spark计算框架以优化处理速度.
主要成果:
- 拟议的算法在数据集中识别了大量的新词 (例如,在微博短消息中16871个).
- 关键词的高热量值得到了实现,这表明了有效的主题识别 (例如,高达0.9991).
- 在基准数据集 (例如,Covid19推文和总统选举推文) 上,在准确性,回忆和F1得分方面展示了强大的绩效指标.
- 在实施Spark框架后,运行速度得到了显著改善.
结论:
- 基于共发生词模型的方法提高了新闻媒体热点挖掘的准确性和效率.
- 该方法为实时主题识别和分析提供了实际意义.
- 集成Spark计算为大规模的文本数据处理提供了可扩展的解决方案.
相关概念视频
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Relative Frequency Histogram
5.4K
The relative frequency depicts the proportion of data points that have each value. The frequency tells the number of data points that have each value. Like the histogram, a relative frequency histogram also has the same shape with a horizontal scale (the x-axis), but the vertical scale (the y-axis) is marked with relative frequencies (percentages of the whole) instead of actual frequencies. A relative frequency histogram is a graphical representation of a frequency distribution where the...
5.4K
Convenience Sampling Method
8.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population.
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
Convenience sampling is a non-random method of sample selection; this method selects individuals that are easily accessible and may result in biased data. For example, a marketing...
8.9K
Contingency Table
2.5K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.5K
Region of Convergence
409
The z-transform is a powerful mathematical tool used in the analysis of discrete-time signals and systems. It is a crucial tool in the analysis of discrete-time systems, but its convergence is limited to specific values of the complex variable z. This range of values, known as the Region of Convergence (ROC), is fundamental in determining the behavior and stability of a system or signal. The ROC defines the region in the complex plane where the z-transform converges, which can take various...
409
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K


