Related Experiment Video
Updated: Jan 13, 2026

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
1.0K
Investigating fine- and coarse-grained structural correspondences between deep neural networks and human object image
Soh Takahashi1, Masaru Sasaki1, Ken Takeda1
1Graduate School of Arts and Science, University of Tokyo, 3-8-1 Komaba, Meguro-ku, 153-8902, Tokyo, Japan.
Summary
Deep neural networks (DNNs) trained with CLIP show human-like object representations at both detailed and broad levels. Self-supervised models capture broad categories but lack fine-grained detail, highlighting language
Area of Science:
- Cognitive Science
- Neuroscience
- Artificial Intelligence
- Computer Vision
Background:
- Understanding how humans form internal object representations is a key challenge in cognitive science.
- Deep neural networks (DNNs) offer a computational framework for studying these mechanisms due to their human-like internal representations.
- Previous research indicates various training methods (supervised, self-supervised, CLIP) yield human-like representations, but fine-grained similarity remains unclear.
Purpose of the Study:
- To investigate whether deep neural network (DNN) object representations match human representations at both coarse and fine-grained levels.
- To compare the representational similarity of models trained with different paradigms (CLIP, self-supervised) to human object representations.
- To elucidate the role of linguistic information in the acquisition of precise object representations.
Main Methods:
- Employed an unsupervised alignment method using Gromov-Wasserstein Optimal Transport to compare human and DNN object representations.
- Assessed fine-grained and coarse-grained matching by estimating optimal mappings between human and model representations for individual objects.
- Utilized human similarity judgments for 1854 objects from the THINGS dataset.
Main Results:
- Models trained with CLIP demonstrated strong fine- and coarse-grained matching with human object representations.
- Self-supervised models exhibited limited matching at both fine- and coarse-grained levels.
- Self-supervised models successfully clustered objects reflecting human coarse category structures.
Conclusions:
- CLIP-trained models capture human object representations more comprehensively, including fine-grained details, compared to self-supervised models.
- Linguistic information, as utilized by CLIP, appears crucial for acquiring precise object representations.
- Self-supervised learning shows potential for capturing coarse categorical structures in object representations.
