COCO:一个对COVID-19阴谋论进行注释的Twitter数据集
Johannes Langguth1,2, Daniel Thilo Schroeder1,3, Petra Filkuková1
1Simula Research Lab, Kristian Augusts Gate 23, Oslo, Norway.
概括
这项研究介绍了3495条关于COVID-19阴谋论的推文的标记数据集. 这些数据有助于训练机器学习模型,用于检测错误信息和分析叙事流行.
科学领域:
- 社交媒体分析 社交媒体分析
- 计算社会科学 计算社会科学
- 公共卫生传播 公共卫生传播
背景情况:
- 随着COVID-19大流行,社交媒体的错误信息显著增加,包括许多阴谋论.
- 这些叙述涵盖了各种主题,并提出了相互竞争的观点,使公众的理解复杂化.
研究的目的:
- 为研究社交媒体上的COVID-19相关阴谋论创建一个全面的数据集.
- 实现机器学习模型的开发,用于错误信息中的立场和主题检测.
- 为了促进阴谋叙事结构和相互联系的定性分析.
主要方法:
- 从2020年1月到2021年6月,收集了3495条与COVID-19相关的推文的数据集.
- 三名专家注释者手动标记了推特对12个阴谋主题的立场,导致近42,000个标签.
- 机器学习模型,特别是BERT,用于立场和主题检测任务.
主要成果:
- 开发的数据集有效地支持机器学习分类器来识别错误信息的立场和主题.
- 定性分析揭示了阴谋叙事及其关系中的共同结构.
- 该研究证明了BERT在联合立场和主题检测方面的成功应用.
结论:
- 创建的数据集是研究COVID-19错误信息和阴谋论的宝贵资源.
- 这些发现提供了关于在线阴谋叙事的性质和传播的见解.
- 这项工作有助于开发用于打击与健康有关的错误信息的工具.
相关概念视频
Causality in Epidemiology
500
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
500
Single Nucleotide Polymorphisms-SNPs
15.3K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.3K
Bias in Epidemiological Studies
375
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
375
Contingency Table
2.5K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.5K
Confounding in Epidemiological Studies
196
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
196
Censoring Survival Data
151
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
151


