Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Confounding in Epidemiological Studies01:27

Confounding in Epidemiological Studies

144
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
144
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

157
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
157
Contingency Table01:29

Contingency Table

2.4K
A contingency table provides a way of portraying data that can facilitate calculating probabilities. It is a method of displaying a frequency distribution as a table with rows and columns to show how two variables may be dependent (contingent) upon each other; The table helps determine conditional probabilities quite quickly and can help systematically organize, analyze and quantify data. The table displays sample values concerning two variables that may be dependent or contingent on one...
2.4K
Multiple Regression01:25

Multiple Regression

2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Aggregates Classification01:29

Aggregates Classification

305
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
305
Classification of Leukocytes01:30

Classification of Leukocytes

1.7K
Leukocytes are classified into two groups based on the presence or absence of cytoplasmic granules. Granular leukocytes, which contain granules, belong to the myeloid lineage and are divided into three subtypes: neutrophils, eosinophils, and basophils. These cells are roughly spherical and characterized by the granules in their cytoplasm.
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
1.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Proactive fault prediction in marine diesel engines using multivariate machine learning.

Scientific reports·2026
Same author

Integrating Deep Learning-Based IoT and Fog Computing with Software-Defined Networking for Detecting Weapons in Video Surveillance Systems.

Sensors (Basel, Switzerland)·2022
Same author

A Reliable and Efficient Tracking System Based on Deep Learning for Monitoring the Spread of COVID-19 in Closed Areas.

International journal of environmental research and public health·2021
Same author

A Novel Interference Avoidance Based on a Distributed Deep Learning Model for 5G-Enabled IoT.

Sensors (Basel, Switzerland)·2021
Same author

Handling varying amounts of missing data when classifying mental-health risk levels.

Studies in health technology and informatics·2014

Related Experiment Video

Updated: Jun 6, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.4K

Enhancing multilabel classification for unbalanced COVID-19 vaccination hesitancy tweets using ensemble learning.

Sherine Nagy Saleh1

  • 1Computer Engineering Department, Arab Academy for Science, Technology and Maritime Transport, Abukir, Alexandria, 1029, Egypt.

Computers in Biology and Medicine
|November 24, 2024
PubMed
Summary

This study enhances vaccine discourse analysis on social media by improving tweet classification. An ensemble BERT model with oversampling significantly boosts performance on imbalanced public health datasets.

Keywords:
BERTClassificationEnsembleMultilabel

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

504
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.5K

Related Experiment Videos

Last Updated: Jun 6, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.4K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

504
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.5K

Area of Science:

  • Social media analytics
  • Public health informatics
  • Computational linguistics

Background:

  • Vaccination is key to public health, but vaccine hesitancy persists globally.
  • Analyzing social media discourse is vital for understanding public sentiment towards vaccines.
  • Class imbalance in social media data presents a significant challenge for accurate analysis.

Purpose of the Study:

  • To improve the classification of tweets related to public health and vaccine discourse.
  • To address the challenge of class imbalance in multilabel, multiclass tweet classification.
  • To evaluate an ensemble-based framework using BERT models for enhanced tweet analysis.

Main Methods:

  • Investigated oversampling techniques to handle unbalanced datasets.
  • Developed an ensemble framework utilizing multiple BERT models fine-tuned on diverse text corpora.
  • Applied the proposed model to a benchmark dataset with 12 imbalanced classes.

Main Results:

  • Achieved significant improvements in performance metrics compared to existing methods.
  • Observed a 2% increase in micro, macro, and weighted average F1 scores.
  • Demonstrated a 4% increase in accuracy and a 3% increase in average Jaccard index.

Conclusions:

  • Fine-tuning BERT models with an ensemble approach and oversampling effectively improves tweet classification on unbalanced public health datasets.
  • The proposed method shows promise for analyzing vaccine discourse on social media.
  • This approach can potentially support public health initiatives by providing better insights into public opinion.