Related Experiment Video
Updated: Jun 20, 2025

Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms
Published on: February 8, 2019
Fine-grained knowledge about manipulable objects is well-predicted by contrastive language image pre-training.
Jon Walbrin1,2, Nikita Sossounov1,2, Morteza Mahdiani3
1Proaction Laboratory, Faculty of Psychology and Educational Sciences, University of Coimbra, Coimbra, Portugal.
Deep learning models, specifically CLIP-ViT, can approximate human object recognition by learning behavioral dimensions from image-text data. This multimodal network excels over image-only models in understanding fine-grained object knowledge.
Area of Science:
- Cognitive Science
- Computer Vision
- Artificial Intelligence
Background:
- Object recognition involves distinguishing similar items, crucial for tasks like tool selection.
- Human object knowledge is organized by behavioral dimensions: vision, manipulation, and function.
- Investigating if deep learning can replicate these human-centric object properties is essential.
Purpose of the Study:
- To determine if deep learning models can approximate the fine-grained behavioral dimensions of human object knowledge.
- To compare the performance of multimodal networks against image-only networks in predicting these dimensions.
Main Methods:
- Utilized CLIP-ViT, a multimodal network trained on extensive image-text pairs.
- Evaluated the model's ability to predict behavioral dimensions relevant to object manipulation and function.
- Compared CLIP-ViT against other networks pre-trained on image-only datasets.
Main Results:
- Behavioral dimensions of object knowledge were generally well-predicted by CLIP-ViT.
- CLIP-ViT demonstrated superior performance compared to networks trained solely on images.
- The findings highlight the model's capacity for approximating nuanced object understanding.
Conclusions:
- CLIP-ViT effectively approximates fine-grained object knowledge, demonstrating the power of multimodal learning.
- Multimodal pre-training and large-scale datasets significantly benefit the model's ability to capture behavioral dimensions.
- This research bridges cognitive science and AI, offering insights into artificial object recognition capabilities.
More Related Videos
Related Concept Videos
Nonconscious Mimicry
Concepts and Prototypes
The brain organizes this information using concepts, which are mental categories grouping linguistic data,...
Observational Learning
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Cognitivism
Previously dominated by behaviorism, which prioritized observable behaviors and largely ignored mental processes, psychology transformed in the 1950s. Cognitive psychologists argue that understanding how we think and process...
Stereotype Content Model

