Related Experiment Video
Updated: Aug 16, 2025

13:19
Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
9.3K
Reddit financial image post sentiment dataset
Alexander Fottner1, Yarema Okhrin1, Jonathan Pfahler1
1Department of Statistics, Faculty of Business and Economics, University of Augsburg, Universitaetsstr. 16, 86159 Augsburg, Germany.
Data in Brief
|December 19, 2022
Summary
This study extracts sentiment from financial subreddit posts, creating a dataset for financial forecasting. The processed data captures market trends and trading decisions from social media discussions.
Area of Science:
- Computational Social Science
- Financial Technology
- Natural Language Processing
Background:
- Financial discussions on social media platforms like Reddit provide insights into market trends and investor sentiment.
- Extracting actionable sentiment data from diverse content formats (text, images) is challenging.
Purpose of the Study:
- To create a novel dataset of sentiment information derived from financial subreddit posts.
- To enable financial forecasting by attributing sentiment to specific financial tickers.
Main Methods:
- Collected data from Reddit and Pushshift APIs, processed via AWS.
- Utilized a fine-tuned MobileNets neural network for image classification into four categories.
- Employed Optical Character Recognition (OCR) and custom methods for text and sentiment extraction from images and text.
Main Results:
- Developed a processed sentiment dataset from financial subreddit image and text posts.
- Classified images into memes, number posts, text posts, and chart posts for targeted analysis.
- Tracked financial tickers to link sentiment to specific financial products.
Conclusions:
- The dataset offers a valuable resource for financial forecasting and analyzing social media sentiment dynamics.
- The methodology allows for the extraction of sentiment signals from complex, multi-modal social media data.
- This work facilitates a deeper understanding of the interplay between social media sentiment and financial markets.
Related Concept Videos
Relative Frequency Histogram
5.6K
The relative frequency depicts the proportion of data points that have each value. The frequency tells the number of data points that have each value. Like the histogram, a relative frequency histogram also has the same shape with a horizontal scale (the x-axis), but the vertical scale (the y-axis) is marked with relative frequencies (percentages of the whole) instead of actual frequencies. A relative frequency histogram is a graphical representation of a frequency distribution where the...
5.6K
Social Proof
27.8K
Social proof is a form of persuasion based on comparison and conformity. People compare their behavior and actions to what others are doing and will change to conform to do what their peers do.
27.8K
Relative Frequency Distribution
11.1K
A relative frequency distribution is the proportion or fraction of times a value occurs in a data set. To find the relative frequencies, one can divide each frequency by the total number of data points in the sample. It is very similar to a regular frequency distribution, except that instead of reporting how many data values fall in a class, a relative frequency distribution reports the fraction of data values that fall in a class. These fractions or proportions are called relative frequencies...
11.1K
Modified Boxplots
9.9K
A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
9.9K
Central Tendency: Analysis
189
Measures of central tendency are tools used in biostatistics to identify the average or center of a dataset. They offer a single representative value for understanding and summarizing data distribution.
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
189
Reaction Quotient
48.9K
The status of a reversible reaction is conveniently assessed by evaluating its reaction quotient (Q). For a reversible reaction described by m A + n B ⇌ x C + y D, the reaction quotient is derived directly from the stoichiometry of the balanced equation as
48.9K

