在物质感知中探究视觉和语言之间的联系,使用心理物理学和无监督学习.
Chenxi Liao1, Masataka Sawayama2, Bei Xiao3
1American University, Department of Neuroscience, Washington DC, United States of America.
PLoS computational biology
|October 3, 2024
概括
人类的视觉和语言在物质感知中显示了适度的相关性. 然而,语言可能会错过视觉细微差别,特别是对于模两可的材料,突出显示了对视觉语言模型的需求.
科学领域:
- 认知科学 认知科学
- 计算机视觉 计算机视觉
- 语言学的语言学.
背景情况:
- 人类擅长视觉材料的歧视,并使用语言来描述它们.
- 了解视觉感知和语义表示之间的联系是人类认知的关键.
研究的目的:
- 研究视觉判断与物质感知中的语言表达之间的关系.
- 探索视觉特征如何映射到语义表示.
- 比较人类对材料的视觉和语言表现.
主要方法:
- 使用深度生成模型生成现实的材料图像,在类别之间创建平稳的过渡.
- 通过视觉材料相似性判断和自由形式的口头描述收集行为数据.
- 采用无监督对齐方法来分析代表性结构.
- 从预先训练的深度神经网络中评估材料表示.
主要成果:
- 在分类层面上发现视觉和语言之间存在中度但显著的相关性.
- 在图像到图像层面发现了结构差异,特别是在模两可的材料中.
- 与口头描述相比,在视觉判断方面观察到更大的个体差异.
- 证明视觉语言模型比仅视觉模型更好地与人类的视觉判断保持一致.
结论:
- 口头描述捕捉粗的材料质量,但可能不能完全代表视觉细微差别.
- 视觉语言关系对于全面的物质感知模型至关重要.
- 建议使用人类行为和计算模型评估交叉模式表示对齐的框架.
相关概念视频
Vision
48.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
48.6K
Depth Perception and Spatial Vision
2.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.7K
Visual System
2.3K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
2.3K
Color Vision
2.0K
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
2.0K
Parallel Processing
961
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
961
Prosopagnosia
1.3K
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
1.3K


