Related Experiment Video
Updated: Aug 16, 2025

03:53
Author Spotlight: Advancements in Multiplex Detection of Respiratory Viruses
Published on: November 10, 2023
1.3K
Multi-label multi-class COVID-19 Arabic Twitter dataset with fine-grained misinformation and situational information
Rasha Obeidat1, Maram Gharaibeh1, Malak Abdullah1
1Department of Computer Science, Jordan University of Science and Technology, Irbid, Jordan.
Peerj. Computer Science
|December 19, 2022
Summary
This study introduces a new Arabic dataset for detecting COVID-19 misinformation and situational information on social media. The research provides fine-grained classification models to combat fake news and improve public health responses.
Area of Science:
- Computational Social Science
- Natural Language Processing
- Public Health Informatics
Background:
- The COVID-19 pandemic saw a surge in social media misinformation, impacting public health and vaccination efforts.
- Existing datasets often lack fine-grained classification for misinformation types and situational context.
- Accurate identification of misinformation sub-types is crucial for targeted interventions and public safety.
Purpose of the Study:
- To develop comprehensive annotation guidelines for fine-grained misinformation and situational information.
- To release the first Arabic COVID-19 misinformation dataset with multi-class and multi-label annotations.
- To establish baseline performance for misinformation and situational information classification models.
Main Methods:
- Developed 19 fine-grained misinformation classes and 6 situational information classes.
- Created an Arabic dataset of approximately 6.7K tweets with multi-class/multi-label annotations.
- Experimented with machine learning and transformer-based classifiers (including AraBERT-COV19) for classification tasks.
Main Results:
- The dataset is validated as suitable for building misinformation and situational information classification models.
- AraBERT-COV19 achieved high performance (81.6% F-score) for multi-class misinformation classification.
- Label Powerset with linear SVC demonstrated strong performance (76.69% F-score) for multi-label misinformation classification.
Conclusions:
- The developed dataset and classification models are vital for combating COVID-19 misinformation in Arabic.
- Fine-grained classification enables more effective strategies for managing online health information.
- This work provides a foundation for future research in Arabic misinformation detection and emergency response.
Keywords:
BERTCOVID-19Data annotationData collectionDeep learningFake newsMachine learningMisinformation detectionSituational informationTransformersMore Related Videos
Related Concept Videos
Classification of Illness
7.7K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.7K
Aggregates Classification
361
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
361
Contingency Table
2.6K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.6K
Classification of Signals
643
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
643
Statistical Methods for Analyzing Epidemiological Data
474
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
474

