Related Experiment Video
Updated: Sep 22, 2025

09:20
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
8.9K
Examining Sentiment in Complex Texts. A Comparison of Different Computational Approaches
Stefan Munnes1, Corinna Harsch1, Marcel Knobloch1
1WZB Berlin Social Science Center, Berlin, Germany.
Frontiers in Big Data
|May 23, 2022
Summary
Computational sentiment analysis of complex texts like German literature reviews shows limitations. While automated methods struggle, a semi-automated approach with human input offers better accuracy for nuanced text analysis.
Area of Science:
- Computational linguistics
- Digital humanities
- Natural language processing
Background:
- Analyzing complex texts like literature reviews presents challenges due to nuanced and ambiguous language.
- Traditional computational sentiment analysis methods may not capture the full spectrum of sentiment in such corpora.
- A metric scale for sentiment analysis is proposed over a dichotomous scale to account for nuanced expressions.
Purpose of the Study:
- To evaluate the accuracy of different computational methods for sentiment analysis of German literature reviews.
- To compare automated and semi-automated approaches against a human-coded "gold standard".
- To provide guidance on selecting appropriate computational methods and pre-processing techniques for complex texts.
Main Methods:
- Comparison of prefabricated dictionaries, self-created dictionaries using word embeddings (pre-trained and self-trained), and a semi-automated approach.
- Sentiment prediction was benchmarked against human-coded sentiments using a metric scale.
- Analysis of word embedding methods, including coding intensity and pre-processing requirements.
Main Results:
- Prefabricated dictionaries showed low to medium correlation (r=0.32-0.39) with human-coded sentiments.
- Self-created dictionaries using word embeddings yielded lower accuracy (r=0.10-0.28).
- A semi-automated approach achieved higher correlation (r≈0.6) but required significant human coding for training.
Conclusions:
- Fully automated computational sentiment analysis is not yet reliable for complex texts like literature reviews.
- Semi-automated approaches, despite requiring human effort, demonstrate greater potential for accurate sentiment analysis in nuanced corpora.
- Researchers should carefully consider method selection and pre-processing levels when analyzing complex texts computationally.
Keywords:
German literatureautomated text analysiscomputer-assisted text analysisdictionaryscaling methodsentiment analysisword embeddingsMore Related Videos
Related Concept Videos
Empathy
9.7K
Some researchers suggest that altruism operates on empathy. Empathy is the capacity to understand another person’s perspective, to feel what he or she feels. An empathetic person makes an emotional connection with others and feels compelled to help (Batson, 1991). Empathy can be expressed in several ways, including cognitive, affective, and motor.
9.7K
Group Design
9.7K
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between...
9.7K
Mass Spectrometry: Complex Analysis
913
Mass spectrometry is an important technique for the identification of pure compounds. However, it has some limitations for the analysis of complex mixtures, often due to excessive fragmentation making the spectrum too complicated to decipher. Mass spectrometry can be combined with suitable separation methods in sequence, forming hyphenated methods, which are useful in the analysis of complex mixtures.
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
913
Multiple Comparison Tests
4.0K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.0K
Central Tendency: Analysis
225
Measures of central tendency are tools used in biostatistics to identify the average or center of a dataset. They offer a single representative value for understanding and summarizing data distribution.
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
The mean is one such measure, calculated by totaling all values in a dataset and dividing by the number of values. For instance, the mean blood pressure reading (120, 130, 140, 150) would be 135. However, the mean can be affected by extreme values or outliers.
The median, another measure,...
225
Stereotype Content Model
14.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.9K

