CVTrack:用于视觉跟踪的组合卷积神经网络和视觉转换器融合模型
Jian Wang1,2, Yueming Song1,2, Ce Song1
1Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, China.
Sensors (Basel, Switzerland)
|January 11, 2024
概括
CVTrack通过将卷积神经网络 (CNN) 和视觉变压器结合在平行双分支网络中来增强对象跟踪. 这种方法有效地融合了本地和全球特征,提高了跟踪精度和实时性能.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 目前的单对象追踪器主要使用卷积神经网络 (CNN) 或视觉转换器 (ViT) 作为其骨干.
- 美国有线电视新闻网在当地特色提取方面表现出色,但在全球代表性方面却扎不堪.
- ViTs通过自我注意力捕捉远程依赖,但可能会错过局部细节.
研究的目的:
- 解决现有的单个物体跟踪方法的局限性.
- 提出一种新的跟踪算法,CVTrack,它集成了本地和全球特征提取能力.
- 为了实现最先进的性能,同时保持实时执行速度.
主要方法:
- 推出了CVTrack,这是一个目标跟踪算法,具有平行双分支骨干网络,结合了CNN和变压器架构.
- 采用双向信息交互通道,有效地融合本地 (CNN) 和全球 (变压器) 功能.
- 使用深度交叉相关和基于变压器的方法,用于全面的模板和搜索区域,在预测之前进行特征融合.
主要成果:
- 在五个基准数据集中,CVTrack实现了最先进的性能.
- 拟议的追踪器展示了实时执行能力.
- 废弃研究证实了平行双分支骨干中的单个模块的有效性.
结论:
- 平行双分支的骨干有效地利用了CNN和变压器的优势,实现了卓越的对象跟踪.
- CVTrack提供了一个强大的解决方案,用于单个对象跟踪,平衡精度和速度.
- 融合策略显著增强模板和搜索区域特征之间的交互.
相关概念视频
Transformers in Distribution System
103
Transformers in distribution systems can be broadly categorized into distribution substation transformers and other distribution transformers. They are crucial for stepping down high transmission voltages to levels suitable for distribution and end-user applications.
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
103
Transformers with Off-Nominal Turns Ratios
157
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
157
Types Of Transformers
977
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
977
Visual System
585
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
585
The Ideal Transformer
395
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
395
Vision
53.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.4K


