Related Experiment Video
Updated: Sep 4, 2025

09:20
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
8.8K
A social and news media benchmark dataset for topic modeling
Samuel Miles1, Lixia Yao2, Weilin Meng2
1Department of Electrical and Computer Engineering, IUPUI, Indianapolis, IN 46202, USA.
Data in Brief
|July 21, 2022
Summary
This study introduces two new datasets for topic modeling research, enabling direct comparison of techniques like pPSO, ETM, and NVDM. These datasets facilitate reproducible research in natural language processing and text analysis.
Area of Science:
- Natural Language Processing
- Machine Learning
- Data Science
Background:
- Topic modeling research faces challenges in comparing techniques due to unavailable data and preprocessing steps.
- Recent advancements focus on vector embeddings with generative and evolutionary topic modeling.
Discussion:
- Presents two secondary datasets derived from a cancer health forum and newsgroups.
- Details preprocessing steps (punctuation, stop word, high frequency word removal).
- Datasets support comparison of pPSO, ETM, and NVDM topic modeling techniques with varying topic numbers (10, 20, 30) and embeddings (sBERT, Skipgram).
Key Insights:
- The datasets enable direct, reproducible comparisons of diverse topic modeling approaches.
- Includes unique identifiers for original documents, preprocessing scripts, and generated topic keywords.
- Provides the algorithm for the pPSO evolutionary topic modeling technique.
Outlook:
- Facilitates standardized evaluation of topic modeling algorithms.
- Encourages further research and development in comparative text analysis.
- Aims to advance the field of unsupervised text representation and understanding.
Related Concept Videos
Cluster Sampling Method
12.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.6K
Social Proof
28.4K
Social proof is a form of persuasion based on comparison and conformity. People compare their behavior and actions to what others are doing and will change to conform to do what their peers do.
28.4K
Stereotype Content Model
14.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.9K
Social Exchange Theory
35.3K
We have discussed why we form relationships, what attracts us to others, and different types of love. But what determines whether we are satisfied with and stay in a relationship? One theory that provides an explanation is social exchange theory. According to social exchange theory, we act as naïve economists in keeping a tally of the ratio of costs and benefits of forming and maintaining a relationship with others (Rusbult & Van Lange, 2003).
35.3K
Social Facilitation
32.7K
Not all intergroup interactions lead to negative outcomes. Sometimes, being in a group situation can improve performance. Social facilitation occurs when an individual performs better when an audience is watching than when the individual performs the behavior alone. This typically occurs when people are performing a task for which they are skilled.
32.7K
Measures of Central Tendency
16.2K
The "center" of a data set is also a way of describing location. The two most widely used measures of the "center" of the data are the mean (average) and the median. The words "mean" and "average" are often used interchangeably. The substitution of one word for the other is common practice. The technical term is "arithmetic mean" and "average" is technically a center location. However, in practice among non-statisticians,...
16.2K

