Investigating fine- and coarse-grained structural correspondences between deep neural networks and human object image

Soh Takahashi1, Masaru Sasaki1, Ken Takeda1

  • 1Graduate School of Arts and Science, University of Tokyo, 3-8-1 Komaba, Meguro-ku, 153-8902, Tokyo, Japan.

Summary

Deep neural networks (DNNs) trained with CLIP show human-like object representations at both detailed and broad levels. Self-supervised models capture broad categories but lack fine-grained detail, highlighting language