弥合视觉和文本语义:实现对无偏见场景图形生成的一致性
概括
场景图形生成 (SGG) 得到了新的视觉文本语义一致性网络 (VTSCN) 的改进. 这种方法将SGG模型作为推理任务,显著减少长尾偏差并增强视觉关系检测.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 认知科学 认知科学
背景情况:
- 场景图形生成 (SGG) 旨在识别图像中的视觉关系.
- 现有的SGG方法存在长尾偏差,限制了它们的实际应用.
- 当前的方法往往过于简化了SGG作为分类任务,阻碍了细粒度细节的捕获,并增加了模两可.
研究的目的:
- 提出一个新的视觉文本语义一致性网络 (VTSCN) 用于场景图形生成.
- 解决并显著缓解SGG中的长尾偏差.
- 将SGG模型作为一个由双过程认知心理学启发的推理过程.
主要方法:
- 引入了一个混合联盟表示 (HUR) 模块,模拟空间意识和工作记忆的快速,自主型1认知过程.
- 开发了一个全球文本语义建模 (GTS) 模块,通过对象对的文本上下文建模来进行高阶推理 (类型2过程).
- 集成了一个异质语义一致性 (HSC) 模块,以平衡1型和2型过程,模仿关联认知.
主要成果:
- 与最先进的方法相比,VTSCN在视觉基因组,GQA和PSG数据集上表现出卓越的性能.
- 废弃研究证实了拟议的VTSCN架构及其组件的有效性.
- VTSCN成功地将SGG模型作为一个推理任务,减轻长尾偏差.
结论:
- 通过结合人类认知过程,VTSCN为SGG模型设计提供了一个新的框架.
- 这种方法有效地减少了长尾偏差,使SGG更实用.
- 这些发现表明,认知心理学原则可以激发更强大,更准确的计算机视觉模型.
更多相关视频
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
7.3K
06:15Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
7.7K
相关概念视频
Schemas
11.6K
A schema is a mental construct consisting of a cluster or collection of related concepts (Bartlett, 1932). There are many different types of schemata, and they all have one thing in common: schemata are a method of organizing information that allows the brain to work more efficiently. When a schema is activated, the brain makes immediate assumptions about the person or object being observed.
11.6K
Visual System
579
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
579
Vector Algebra: Graphical Method
12.1K
Vectors can be multiplied by scalars, added to other vectors, or subtracted from other vectors. The vector sum of two (or more) vectors is called the resultant vector or, for short, the resultant.
We use the laws of geometry to construct resultant vectors, followed by trigonometry to find vector magnitudes and directions. For a geometric construction of the sum of two vectors in a plane, we follow the parallelogram rule. Suppose two vectors are at arbitrary positions. Translate either one of...
We use the laws of geometry to construct resultant vectors, followed by trigonometry to find vector magnitudes and directions. For a geometric construction of the sum of two vectors in a plane, we follow the parallelogram rule. Suppose two vectors are at arbitrary positions. Translate either one of...
12.1K
Gestalt Principles of Perception
298
Gestalt principles provide a framework for understanding how humans perceive objects as unified wholes within their context. These principles are essential in explaining the cognitive processes that make sense of complex visual stimuli by organizing them into coherent groups. One fundamental principle is proximity, which posits that objects located close to each other are perceived as a collective group. For instance, when dots are positioned near one another, the visual system interprets them...
298
Spanning Openings in Brick Walls
187
In brick wall construction, supporting structures are crucial for openings like windows and doors to maintain the integrity and support the weight of the wall above. These supports include lintels, corbels, and arches, each serving specific structural purposes.
Lintels are primary supports used to span openings and can be crafted from materials such as reinforced concrete, steel-reinforced brick masonry, or simple steel angles. These are straightforward to install and are typically concealed...
Lintels are primary supports used to span openings and can be crafted from materials such as reinforced concrete, steel-reinforced brick masonry, or simple steel angles. These are straightforward to install and are typically concealed...
187
Depth Perception and Spatial Vision
643
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
643
