Related Experiment Video
Updated: Sep 23, 2025

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
606
An iterative topic model filtering framework for short and noisy user-generated data: analyzing conspiracy theories
Gillian Kant1, Levin Wiebelt1, Christoph Weisser2
1University of Göttingen, Göttingen, Germany.
Summary
This study introduces Iterative Filtering and Hashtag Pooling to analyze Twitter data on conspiracy theories. During late 2020, "Election Fraud" and "Covid-19-hoax" theories were prominent, linked to US election and pandemic discussions.
Area of Science:
- Social Media Analysis
- Computational Social Science
- Natural Language Processing
Background:
- Conspiracy theories are increasingly prevalent, amplified by social media and influencing public opinion.
- Analyzing online discourse, particularly on platforms like Twitter, is crucial for understanding societal trends.
- Twitter data presents challenges due to its noisy and sparse nature, requiring advanced pre-processing techniques.
Purpose of the Study:
- To introduce and evaluate a novel pre-processing approach, Iterative Filtering with Hashtag Pooling, for analyzing Twitter data.
- To investigate the dynamics of conspiracy theories on US Twitter during the final four months of 2020.
- To identify and characterize the main conspiracy-related topics and their association with major events like the US Presidential Elections and Covid-19.
Main Methods:
- Development of Iterative Filtering for Latent Dirichlet Allocation (LDA) pre-processing.
- Implementation of Hashtag Pooling as an additional data refinement step.
- Application of geo-spatial Twitter data and LDA Topic Models to monitor US public discourse.
Main Results:
- The Iterative Filtering and Hashtag Pooling approach effectively handles noisy Twitter data for topic modeling.
- During the study period (late 2020), mainstream conspiracy theories were overshadowed by US Presidential Election and Covid-19 discussions.
- The primary conspiracy theories identified were related to "Election Fraud" and the "Covid-19-hoax," often co-occurring with Trump-related keywords.
Conclusions:
- The proposed methods enhance the ability to study online discourses and identify relevant tweets.
- Conspiracy theories related to the 2020 US election and Covid-19 were significant but secondary to broader discussions on these events.
- The findings highlight the interconnectedness of political discourse, public health events, and the spread of specific conspiracy narratives online.
Keywords:
Conspiracy theoriesCovid-19Geo-spatial analysisHashtag poolingIterative filteringLDALatent Dirichlet allocationNLPNLP pre-processingSARS-CoV-2Sentiment analysisMore Related Videos
Related Concept Videos
Group Polarization
35.8K
Group polarization is the strengthening of an original group attitude following the discussion of views within a group (Teger & Pruitt, 1967). That is, if a group initially favors a viewpoint, after discussion the group consensus is likely a stronger endorsement of the viewpoint. Conversely, if the group was initially opposed to a viewpoint, group discussion would likely lead to stronger opposition.
35.8K
Filtration
1.0K
Filtration is a physical separation process that involves passing a suspension through a porous medium to separate solids from fluids. During filtration, solids collect on the porous medium while liquids, also collectively known as the filtrate, pass through. The filtration medium is selected based on the filtration purpose, quantity, and nature of the precipitate. The general criteria for a suitable filtering medium are that it is inert, mechanically strong, nonabsorbent toward dissolved...
1.0K
Sampling Theorem
813
In signal processing, the analysis of continuous-time signals, denoted as x(t), often involves sampling techniques to convert these signals into discrete-time signals. This process is essential for digital representation and manipulation. A critical component in sampling is the train of impulses, characterized by the sampling interval and the sampling frequency. The relationship between these parameters and the original signal's properties dictates the success of the sampling process.
813
Outliers and Influential Points
4.3K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.3K
Null and Alternative Hypotheses
10.3K
The actual hypothesis testing begins by considering two hypotheses. They are termed the null hypothesis and the alternative hypothesis. These hypotheses contain opposing viewpoints.
The null hypothesis, denoted by H0 is a statement of no difference between the variables—they are not related. This can often be considered the status quo. As a result if you cannot accept the null, it requires some action.
The alternative hypothesis, denoted by H1 or Ha, is a claim about the...
The null hypothesis, denoted by H0 is a statement of no difference between the variables—they are not related. This can often be considered the status quo. As a result if you cannot accept the null, it requires some action.
The alternative hypothesis, denoted by H1 or Ha, is a claim about the...
10.3K
Statistical Hypothesis Testing
2.1K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
2.1K

