Related Experiment Video
Updated: Nov 1, 2025

07:31
Investigating the Effect of Visual Imagery and Learning Shape-Audio Regularities on Bouba and Kiki
Published on: September 13, 2019
10.3K
How quantifying the shape of stories predicts their success
Olivier Toubia1, Jonah Berger2, Jehoshua Eliashberg2
1Marketing Division, Columbia Business School, Columbia University, New York, NY 10027; ot2107@gsb.columbia.edu.
Summary
This study quantifies discourse movement in semantic space, revealing how narrative path features influence cultural success across movies, TV, and academic papers. These findings offer insights into popularity and discourse analysis.
Area of Science:
- Computational Linguistics
- Natural Language Processing
- Cultural Analytics
Background:
- Discourse is often described qualitatively (e.g., fast-paced, meandering).
- Quantifying discourse movement and its impact on success remains underexplored.
- Understanding discourse dynamics can illuminate cultural phenomena.
Purpose of the Study:
- To develop and apply quantitative measures for discourse semantic paths.
- To investigate the relationship between semantic path features and text success.
- To explore cross-domain differences in discourse and success metrics.
Main Methods:
- Utilizing advanced natural language processing (NLP) and machine learning (ML).
- Representing texts as sequences in a high-dimensional latent semantic space.
- Developing and applying novel measures to quantify semantic path characteristics.
Main Results:
- Identified quantifiable features of semantic paths in diverse texts.
- Demonstrated a link between semantic path features and measures of success (e.g., citations).
- Observed significant cross-domain variations in discourse patterns and their impact.
Conclusions:
- A generalizable framework for analyzing discourse movement in semantic space has been established.
- The study provides insights into the drivers of cultural success and popularity.
- NLP and ML offer powerful tools for understanding discourse and cultural impact.
Related Concept Videos
Survival Curves
412
Survival curves are graphical representations that depict the survival experience of a population over time, offering an intuitive way to track the proportion of individuals who remain event-free at each time point. These curves are widely used in fields such as medicine, public health, and reliability engineering to visualize and compare survival probabilities across different groups or conditions.
The Kaplan-Meier estimator is the most common method for constructing survival curves. This...
The Kaplan-Meier estimator is the most common method for constructing survival curves. This...
412
Unusual Results
3.5K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
3.5K
Kaplan-Meier Approach
338
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
338
5-Number Summary
5.2K
In a dataset, the 5-number summary includes the minimum data value, the data value of the first quartile, the median data value or data value of the second quartile, the data value of the third quartile, and the maximum data value. These 5 data values can be visualized as a box and whisker plot.
In a box plot, the minimum and maximum data values represent the lower and upper whiskers in the graph, and the median is designated as the center of the box in the chart. The first quartile and third...
In a box plot, the minimum and maximum data values represent the lower and upper whiskers in the graph, and the median is designated as the center of the box in the chart. The first quartile and third...
5.2K
Probability Histograms
12.6K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
12.6K
Hindsight Biases
4.1K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
4.1K

