Comparable 2022 General Election Advertising Datasets from Meta and Google.
Meiqing Zhang1, Furkan Cakmak1, Markus Neumann2
1Wesleyan University, Wesleyan Media Project, Middletown, 06459, USA.
Scientific Data
|June 9, 2025
Summary
This study presents datasets on 2022 U.S. midterm election digital ads from Meta and Google. The data includes processed audiovisual and textual information for analyzing political campaign strategies and public engagement.
Area of Science:
- Political Science
- Communication Studies
- Data Science
Background:
- Digital advertising significantly influences U.S. federal elections.
- Understanding online political campaigns requires accessible, comprehensive data.
Purpose of the Study:
- To introduce two novel datasets detailing digital ads from Meta and Google during the 2022 U.S. midterm elections.
- To provide researchers with structured data for analyzing digital political advertising.
Main Methods:
- Data collection via ad transparency libraries and web scraping from Meta (Facebook, Instagram) and Google (YouTube).
- Processing included automatic speech recognition (ASR), face recognition, and optical character recognition (OCR).
- Data enrichment through labeling and classification tasks for enhanced comparability and utility.
Main Results:
- Creation of two comprehensive datasets on U.S. federal election digital ads (2022 midterm).
- Datasets contain metadata, ASR transcripts, OCR text, face recognition data, and classification labels.
- The processed data enables detailed analysis of ad content and campaign strategies.
Conclusions:
- The datasets offer valuable resources for analyzing the digital election advertising landscape.
- Facilitates research in political science, communication, and data science for insights into digital political campaigns.
- Enables deeper understanding of campaign strategies and public engagement in online environments.
Related Concept Videos
Factorial Design
13.0K
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
13.0K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Margin of Error
4.0K
The margin of error is also called the maximum error of an estimate. The margin of error is the maximum possible or expected difference between the observed sample parameter value and the actual population parameter value. For proportion, it is the maximum difference between the value of sample proportion obtained from the data and the true value of population proportion. As the true value of the population parameter is not known, the margin of error is calculated using the sample statistic.
4.0K
Data Collection by Experiments
24.0K
Data collection is a systematic method of obtaining, observing, measuring, and analyzing accurate information. An experimental study is a standard method of data collection that involves the manipulation of the samples by applying some form of treatment prior to data collection. It refers to manipulating one variable to determine its changes on another variable. The sample subjected to treatment is known as “experimental units.”
An example of the experimental method is a public...
An example of the experimental method is a public...
24.0K
Geometric Mean
3.4K
The mean is a measure of the central tendency of a data set. In some data sets, the data is inherently multiplicative, and the arithmetic mean is not useful. For example, the human population multiplies with time, and so does the credit amount of financial investment, as the interest compounds over successive time intervals.
In cases of multiplicative data, the geometric mean is used for statistical analysis. First, the product of all the elements is taken. Then, if there are n elements in the...
In cases of multiplicative data, the geometric mean is used for statistical analysis. First, the product of all the elements is taken. Then, if there are n elements in the...
3.4K


