基于图形的视觉转换器,可用于从头开始对小数据集进行培训
Peng Li1, Lu Huang2,3, Jin Li2
1Emergency Department, Yantaishan Hospital, Yantai, Shandong, 264008, China.
Scientific reports
|July 8, 2025
概括
基于图形的视觉转换器 (GvTs) 通过结合诱导偏差来弥合小型数据集的性能差距. 在没有预先培训的情况下,GvT的表现优于标准的视觉转换器 (ViT).
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 视觉转换器 (ViT) 在大规模图像分类方面表现出色,但在较小的数据集上,由于缺乏诱导偏差,与卷积神经网络 (CNN) 相比,其表现不佳.
- 这种性能差异凸显了对架构修改的需求,以在数据有限的场景中提高ViT效率.
研究的目的:
- 引入基于图形的视觉转换器 (GvT),旨在通过集成基于图形的机制来提高小数据集的ViT性能.
- 解决标准ViT在捕获局部空间信息方面的局限性,并缓解注意力机制中的低级瓶.
主要方法:
- 拟议的GvT采用了对查询和密钥的图形卷积投影,在每个块内使用空间邻矩阵.
- 图形卷积用于生成值,并应用说话头技术来克服注意力头的低级瓶.
- 在中间块之间集成了图形聚合,以减少令牌数量并增强语义信息聚合.
主要成果:
- 在小数据集上,GvT表现出与深度CNN相美或更高的性能.
- GvT模型超过了标准ViT的性能,这些ViT没有在大型数据集上进行预训练.
结论:
- GvT架构有效地引入了诱导偏差,显著提高了视觉转换器在小数据集上的性能.
- 在训练数据有限的情况下,GvT为传统的CNN和标准ViT提供了一个有希望的替代方案.
相关概念视频
Depth Perception and Spatial Vision
952
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
952
Vision
55.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.4K
Vector Algebra: Graphical Method
14.0K
Vectors can be multiplied by scalars, added to other vectors, or subtracted from other vectors. The vector sum of two (or more) vectors is called the resultant vector or, for short, the resultant.
We use the laws of geometry to construct resultant vectors, followed by trigonometry to find vector magnitudes and directions. For a geometric construction of the sum of two vectors in a plane, we follow the parallelogram rule. Suppose two vectors are at arbitrary positions. Translate either one of...
We use the laws of geometry to construct resultant vectors, followed by trigonometry to find vector magnitudes and directions. For a geometric construction of the sum of two vectors in a plane, we follow the parallelogram rule. Suppose two vectors are at arbitrary positions. Translate either one of...
14.0K
Improving Translational Accuracy
2.7K
2.7K
Survival Tree
166
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
166
Transformers with Off-Nominal Turns Ratios
213
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
213


