视觉空间关系敏感的变压器用于图像标题
Xianghua Piao1,2, Dong Jin1,2, Min Jung Kwon3
1Department of Computer Science and Engineering, Sejong University, Seoul, 05006, South Korea.
Scientific reports
|December 24, 2025
概括
这项研究引入了一个新的视觉空间关系敏感变压器 (VRST) 用于图像标题. 它通过对准网格和区域特征来改善视觉语义理解,增强关系和空间细节以获得更好的描述.
科学领域:
- 计算机视觉 计算机视觉
- 自然语言处理自然语言处理.
- 人工智能的人工智能
背景情况:
- 图像标题集成了计算机视觉和NLP来描述视觉内容.
- 当前的方法在结合网格和区域特征时,与空间错位作斗争,阻碍了视觉语义连贯性.
- 准确的关系和上下文信息对于生成有意义的图像描述至关重要.
研究的目的:
- 通过解决空间错位问题,提出一种新的方法来增强图像标题.
- 通过对齐多层次的特征来开发统一的视觉空间表示.
- 提高模型捕捉全球关系线索和局部空间细节的能力.
主要方法:
- 引入了空间对齐位置编码器 (SAPE) 来编码对齐的网格级和区域级特征.
- 开发了集团规范化多头注意力 (GNMA) 用于全球关系线索和基于卷积的特征增强注意力 (CFEA) 用于本地细节.
- 提出可学习的自适应位置编码器 (LAPE),以在深度训练期间保存位置信号.
- 将这些组件集成到基于变压器的架构中,称为视觉空间关系敏感变压器 (VRST).
主要成果:
- 在MSCOCO数据集上获得了141.9的CIDEr分数 (Karpathy测试分割).
- 在MSCOCO官方评估服务器上获得了138.2的CIDEr评分.
- 与几个强大的基线模型相比,表现出优越的性能.
结论:
- 拟议的VRST有效地解决了图像标题中的空间错位问题.
- 整合SAPE,GNMA,CFEA和LAPE显著增强了视觉语义关系建模.
- 该方法在MSCOCO数据集上取得了最先进的结果,展示了其有效性.
相关概念视频
Types Of Transformers
1.4K
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
1.4K
Transformers
1.7K
A device that transforms voltages from one value to another using induction is called a transformer. A transformer consists of two separate coils, or windings, wrapped around the same soft iron core. However, they are electrically insulated from each other.
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
1.7K
Transformers with Off-Nominal Turns Ratios
490
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
490
The Ideal Transformer
1.3K
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's tangential...
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's tangential...
1.3K
Visual System
1.6K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.6K
Depth Perception and Spatial Vision
1.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.7K


