When 'Yes' Meets 'But': Can AI Comprehend Contradictory Humor in Comics?
Summary
Large vision language models (VLMs) struggle with humor, especially contradictory comics. New benchmark YESBUT reveals significant performance gaps compared to humans, highlighting needs for better reasoning and cultural understanding in AI.
Area of Science:
- Artificial Intelligence
- Cognitive Science
- Computational Linguistics
Background:
- Large vision language models (VLMs) face challenges in understanding complex humor requiring comparative reasoning.
- This limitation impacts AI's capacity for human-like reasoning and cultural expression.
Purpose of the Study:
- To analyze humor in comics that use contradictory juxtapositions.
- To introduce and utilize the YESBUT benchmark for evaluating VLMs' narrative and comparative reasoning abilities.
- To identify specific weaknesses in current VLM performance.
Main Methods:
- Development of the YESBUT benchmark (1,262 comic images) with diverse contexts and annotations.
- Systematic evaluation of various VLMs on four tasks, focusing on comparative reasoning.
- Investigation of text-based training and social knowledge augmentation strategies.
Main Results:
- Advanced VLMs significantly underperform compared to human capabilities in understanding comic humor.
- Common failure points include visual perception, key element identification, comparative analysis, and hallucinations.
- Current models lack deep narrative understanding, especially with contradictory elements.
Conclusions:
- VLMs exhibit critical weaknesses in understanding cultural and creative expressions, particularly humor based on contradiction.
- The YESBUT benchmark provides a pathway for assessing and improving VLM narrative reasoning.
- Future research should focus on developing context-aware models with enhanced comparative reasoning skills.
More Related Videos
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
10.3K
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
16.6K
Related Concept Videos
Counterfactual Thinking
436
Counterfactual thinking is a cognitive process wherein individuals mentally reconstruct alternative versions of past events, often beginning with “what if” or “if only.” This reflective mechanism plays a significant role in shaping emotional experiences and guiding future behavior. Though typically triggered by unfavorable or unexpected outcomes, counterfactual thinking can also emerge in mundane, everyday decisions and experiences, revealing its deep entrenchment in...
436
Cognitive Dissonance
29.1K
Social psychologists have documented that feeling good about ourselves and maintaining positive self-esteem is a powerful motivator of human behavior (Tavris & Aronson, 2008). In the United States, members of the predominant culture typically think very highly of themselves and view themselves as good people who are above average on many desirable traits (Ehrlinger, Gilovich, & Ross, 2005). Often, our behavior, attitudes, and beliefs are affected when we experience a threat to our...
29.1K
Hypothesis: Accept or Fail to Reject?
28.9K
The outcome of any hypothesis testing leads to rejecting or not rejecting the null hypothesis. This decision is taken based on the analysis of the data, an appropriate test statistic, an appropriate confidence level, the critical values, and P-values. However, when the evidence suggests that the null hypothesis cannot be rejected, is it right to say, 'Accept' the null hypothesis?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
28.9K
Attitudes
25.9K
Attitude is our evaluation of a person, an idea, or an object. We have attitudes for many things ranging from products that we might pick up in the supermarket to people around the world to political policies. Typically, attitudes are favorable or unfavorable: positive or negative (Eagly & Chaiken, 1993). And, they have three components: an affective component (feelings), a behavioral component (the effect of the attitude on behavior), and a cognitive component (belief and knowledge;...
25.9K
Social Proof
24.9K
Social proof is a form of persuasion based on comparison and conformity. People compare their behavior and actions to what others are doing and will change to conform to do what their peers do.
24.9K
Inductive Reasoning
59.0K
Inductive reasoning is a form of logical thinking that uses related observations to arrive at a general conclusion. It is uncertain and operates in degrees to which the conclusions are credible. As such, inductive arguments can be weak or strong, rather than valid or invalid, and conclusions can be used to formulate testable, falsifiable hypotheses.
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
59.0K
