An Effective BERT-Based Pipeline for Twitter Sentiment Analysis: A Case Study in Italian
Marco Pota1, Mirko Ventura1, Rosario Catelli1,2
1Institute for High Performance Computing and Networking (ICAR), National Research Council, 80131 Naples, Italy.
This study presents a novel approach to sentiment analysis for tweets, converting jargon into plain text before analysis with BERT. This method enhances accuracy and language adaptability for better social media insights.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Social Media Analysis
Background:
- Growing interest in sentiment analysis for social media, particularly tweets.
- Current state-of-the-art models often require training from scratch on tweet-specific data.
- Challenges in handling Twitter jargon, including emojis and emoticons.
Purpose of the Study:
- To introduce a novel, two-step approach for Twitter sentiment analysis.
- To improve sentiment analysis by transforming tweet jargon into plain text.
- To leverage pre-trained language models on plain text for enhanced performance and language independence.
Main Methods:
- A two-step process: 1. Transforming tweet jargon (emojis, emoticons) into language-independent plain text. 2. Classifying transformed tweets using BERT pre-trained on plain text corpora.
- Utilizing language-independent procedures for jargon transformation.
- Employing BERT models pre-trained on extensive plain text datasets.
Main Results:
- Demonstrated effectiveness of the proposed approach in a case study on Italian sentiment analysis.
- Achieved competitive or superior performance compared to existing Italian solutions.
- The approach showed promise for application in other languages due to its general methodology.
Conclusions:
- The proposed method offers an effective and adaptable solution for Twitter sentiment analysis.
- Leveraging pre-trained BERT models on plain text is advantageous over training from scratch on tweet-only data.
- The approach's language-independent nature and reliance on readily available pre-trained models make it a promising tool for multilingual sentiment analysis.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
Related Concept Videos
SBAR II: Application of SBAR
SBAR Report from a Nurse to a Health Care Provider
S: "Hello, Dr. Smith. This is Jane, RN, from the Med Surg unit. I am calling to tell you about Ms. White in Room 210, who is experiencing increased pain and redness at her incision site. Her recent...
Automatic Processing and Automatic Social Behavior
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Improving Translational Accuracy
Improving Translational Accuracy
