Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

Spatial-Niche Perspective on the Heterogeneity and Functional Reprogramming of Tumor-Associated Macrophages in Digestive System Tumors.

Cells·2026
Same author

Three-Dimensional Facial Aesthetic Analysis of Oval and Rectangular Face Shapes in Han Chinese Women for Plastic Surgery Applications.

The Journal of craniofacial surgery·2026
Same author

Application of digital health technologies in hypertension self-management: a narrative review.

Frontiers in public health·2026
Same author

Unveiling VARS1: a key driver of colorectal cancer progression and immune modulation.

International journal of clinical oncology·2026
Same author

Belonging in smaller spaces: student voices reshaping inclusion.

BMC psychology·2026
Same author

Inflammation-Driven JNK Activation Promotes EMT and Metastasis in Gastric Cancer and Is Attenuated by Huangjin Shuangshen Granules.

Pharmaceuticals (Basel, Switzerland)·2026

相关实验视频

Updated: Jan 10, 2026

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

727

RepAttn3D:通过时空增强来重新参数化3D注意力,以获得视频理解.

Xiusheng Lu1, Lechao Cheng2, Sicheng Zhao3

  • 1School of Software, Tsinghua University, Beijing, 100084, China.

Neural networks : the official journal of the International Neural Network Society
|November 23, 2025
PubMed
概括

这项研究引入了一个新的空间时间增强3D注意力 (STA-3DA) 模块,以提高视频理解. 该方法增强了功能学习,并降低了用于视频分析的变压器模型的计算成本.

关键词:
三维注意力 3D注意力行动认可 行动认可重新参数化的重新参数化空间时空的连贯性之前的先行性.

更多相关视频

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1000
Photorealistic Learned Landscapes for Augmented Reality
06:54

Photorealistic Learned Landscapes for Augmented Reality

Published on: June 27, 2025

651

相关实验视频

Last Updated: Jan 10, 2026

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

727
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1000
Photorealistic Learned Landscapes for Augmented Reality
06:54

Photorealistic Learned Landscapes for Augmented Reality

Published on: June 27, 2025

651

科学领域:

  • 计算机视觉 计算机视觉
  • 人工智能的人工智能
  • 机器学习 机器学习

背景情况:

  • 结构重新参数化在使用CNN和MLP的图像任务中很常见.
  • 在视频分析中,重新参数化与注意力机制的整合尚未得到充分研究.
  • 视频分析面临着高的计算成本,特别是在推理过程中.

研究的目的:

  • 研究用于视频理解的3D注意力机制的重新参数化.
  • 在增强视频功能学习之前,要整合一个时空连贯性.
  • 解决视频分析任务中的计算挑战.

主要方法:

  • 为变压器架构提出一个空间时间增强的3D注意力 (STA-3DA) 模块.
  • 在培训期间整合3D,空间和时间注意力分支.
  • 将注意力分支合并为一个单一的3D操作,在测试过程中学习重量.

主要成果:

  • 在STA-3DA模块学习更强大的视频功能与微不足道的推断开销.
  • 拟议的模块有效地取代了变压器模型中的标准3D注意力,提高了性能.
  • 在Kinetics-400和Something-Something V2数据集上实现了竞争性视频理解性能.

结论:

  • STA-3DA模块提供了一种高效有效的方法来提高视频理解.
  • 用时空先验对3D注意力进行重新参数化是视频分析的一个有希望的方向.
  • 该方法为减少视频变压器模型中的计算成本提供了实际解决方案.