RTS-ViT:实时共享视觉转换器用于图像分类.
IEEE journal of biomedical and health informatics
|March 3, 2025
概括
一个新的双分支视觉变压器通过融合来自不同补丁大小的特征来增强视网膜图像分类. 这种方法可以在较少的计算资源和没有预先培训的情况下实现卓越的性能,从而提高自我学习能力.
科学领域:
- 计算机视觉 计算机视觉
- 医疗成像医学成像
- 人工智能的人工智能
背景情况:
- 视觉变压器 (ViTs) 在图像分类方面出色.
- 双分支ViT通过融合改善了功能生成.
- 视网膜图像分类需要强大的特征表示.
研究的目的:
- 开发一个高效的双分支视觉转换器用于视网膜图像分类.
- 增强ViT中的特征学习和融合机制.
- 为了提高分类准确性,同时减少计算复杂性.
主要方法:
- 提出了一种带有实时共享功能编码器的双分支视觉变压器.
- 通过独立的分支处理了基础和大尺寸的图像补丁.
- 使用L-Times Attention Fusion (矢量连接和元素智能加法) 实现了多阶段的特征融合.
主要成果:
- 与Cross-ViT相比,实现了更高的性能,平均TOP-1精度高出5.61%.
- 证明了较低的FLOP和较少的模型参数.
- 展示了强大的自我学习能力,而不依赖预先训练的重量.
结论:
- 拟议的双分支ViT与实时共享功能显著提高了视网膜图像分类.
- L-Times注意力融合方法提供了高效和有效的功能集成.
- 该方法为医学图像分析提供了一个有希望的,资源高效的解决方案.
相关概念视频
Transformers in Distribution System
98
Transformers in distribution systems can be broadly categorized into distribution substation transformers and other distribution transformers. They are crucial for stepping down high transmission voltages to levels suitable for distribution and end-user applications.
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
98
Types Of Transformers
943
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
943
Transformers with Off-Nominal Turns Ratios
129
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
129
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K
The Ideal Transformer
342
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
342
Depth Perception and Spatial Vision
508
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
508


