Related Experiment Video
Updated: Jul 7, 2025

Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms
Published on: February 8, 2019
Interpretability Is in the Mind of the Beholder: A Causal Framework for Human-Interpretable Representation Learning.
Emanuele Marconato1,2, Andrea Passerini1, Stefano Teso1,3
1Dipartimento di Ingegneria e Scienza dell'Informazione, University of Trento, 38123 Trento, Italy.
This research introduces a mathematical framework for human-interpretable representation learning (hrl) to create AI explanations understandable by humans. It models human understanding to align AI concepts with human vocabulary, improving AI interpretability.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Human-Computer Interaction
Background:
- Explainable AI (XAI) research is shifting towards concept-based explanations.
- Current methods for acquiring interpretable concepts lack standardization and neglect human understanding.
- A key challenge is modeling the human element in representation learning.
Purpose of the Study:
- To propose a mathematical framework for acquiring interpretable representations.
- To bridge the gap between human and algorithmic interpretability in AI.
- To establish a foundation for future research in human-interpretable representation learning.
Main Methods:
- Developed a formalization of human-interpretable representation learning (hrl).
- Integrated causal representation learning principles.
- Modeled a human stakeholder as an external observer to define alignment.
Main Results:
- Derived a principled notion of alignment between machine representations and human concept vocabularies.
- Linked alignment and interpretability via a name transfer game.
- Clarified relationships between alignment, disentanglement, concept leakage, and content-style separation.
Conclusions:
- The proposed framework offers a principled approach to human-interpretable representation learning.
- Alignment is a crucial factor for creating AI representations understandable by humans.
- This work provides a stepping stone for advancing AI interpretability research.
Related Concept Videos
The Representativeness Heuristic
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Deductive Reasoning
For example, a researcher can deduce specific predictions...
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Cognitivism
Previously dominated by behaviorism, which prioritized observable behaviors and largely ignored mental processes, psychology transformed in the 1950s. Cognitive psychologists argue that understanding how we think and process...
Purposive Learning

