通过层次偏移补偿和UAV空中图像的时间内存更新进行视频实例细分
Ying Huang1, Yinhui Zhang1, Zifen He1
1Faculty of Mechanical and Electrical Engineering, Kunming University of Science and Technology, Kunming 650500, China.
Sensors (Basel, Switzerland)
|July 30, 2025
概括
本研究引入了一种使用无人机 (UAV) 进行视频实例分割 (VIS) 的新方法,通过增强特征捕获和时间建模来提高对变形目标的准确性.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
背景情况:
- 现有的视频实例细分 (VIS) 方法无法准确地细分无人机 (UAV) 镜头中的变形目标.
- 挑战包括无效的特征偏移捕获和不充分的时间相关性建模,导致结果不一致.
研究的目的:
- 为视频实例分割 (HT-VIS) 提出一个具有高概括能力的层次偏移补偿和时间内存更新方法.
- 提高无人机应用中不规则变形目标的VIS的准确性和稳定性.
主要方法:
- 开发了一种层次偏移补偿 (HOC) 模块,用于跨的可变形偏移,顺序和并行捕获空间运动特征.
- 实现了一个时间内存更新 (TMU) 模块,使用卷积长短期内存 (ConvLSTM) 来建模时间动态上下文和更新功能.
主要成果:
- 拟议的HT-VIS方法在YouTubeVIS-2019和自建UAV-Seg数据集上表现出卓越的性能.
- 取得了最先进的结果,在特定数据集上超过CrossVIS高达3.9%和SipMask高达2.1%.
结论:
- HT-VIS框架有效地解决了在UAV智能检查任务中对变形目标进行细分的局限性.
- 该方法显示了平均细分精度的显著改进,并在各种数据集中展示了稳定性.
相关概念视频
Relative Motion Analysis using Rotating Axes
536
Consider a component AB undergoing a linear motion. Along with a linear motion, point B also rotates around point A. To comprehend this complex movement, position vectors for both points A and B are established using a stationary reference frame.
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
536
Relative Motion Analysis using Rotating Axes-Problem Solving
451
Consider a crane whose telescopic boom rotates with an angular velocity of 0.04 rad/s and angular acceleration of 0.02 rad/s2. Along with the rotation, the boom also extends linearly with a uniform speed of 5 m/s. The extension of the boom is measured at point D, which is measured with respect to the fixed point C on the other end of the boom. For the given instant, the distance between points C and D is 60 meters.
Here, in order to determine the magnitude of velocity and acceleration for point...
Here, in order to determine the magnitude of velocity and acceleration for point...
451
Depth Perception and Spatial Vision
929
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
929
Relative Motion Analysis - Velocity
436
A stroke engine has a slider-crank mechanism that converts rotational motion from the crank into linear motion of the slider or vice versa. This mechanism consists of three main parts: the crank, the connecting rod, and the slider.
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
436


