Related Experiment Video
Updated: Jul 25, 2025

Quantification and Whole Genome Characterization of SARS-CoV-2 RNA in Wastewater and Air Samples
Published on: June 30, 2023
COCO: an annotated Twitter dataset of COVID-19 conspiracy theories
Johannes Langguth1,2, Daniel Thilo Schroeder1,3, Petra Filkuková1
1Simula Research Lab, Kristian Augusts Gate 23, Oslo, Norway.
Abstract:
The COVID-19 pandemic has been accompanied by a surge of misinformation on social media which covered a wide range of different topics and contained many competing narratives, including conspiracy theories. To study such conspiracy theories, we created a dataset of 3495 tweets with manual labeling of the stance of each tweet w.r.t. 12 different conspiracy topics. The dataset thus contains almost 42,000 labels, each of which determined by majority among three expert annotators. The dataset was selected from COVID-19 related Twitter data spanning from January 2020 to June 2021 using a list of 54 keywords. The dataset can be used to train machine learning based classifiers for both stance and topic detection, either individually or simultaneously. BERT was used successfully for the combined task. The dataset can also be used to further study the prevalence of different conspiracy narratives. To this end we qualitatively analyze the tweets, discussing the structure of conspiracy narratives that are frequently found in the dataset. Furthermore, we illustrate the interconnection between the conspiracy categories as well as the keywords.
Related Concept Videos
Causality in Epidemiology
Single Nucleotide Polymorphisms-SNPs
Bias in Epidemiological Studies
Contingency Table
Confounding in Epidemiological Studies
Censoring Survival Data

