Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Parallel Processing01:20

Parallel Processing

145
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
145
Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

600
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
600

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same journal

Trap tales: The influence of red alder stand conditions and forest fragmentation on family-level beetle bycatch diversity.

PloS one·2026
Same journal

MamNet-PT: A Mamba-enhanced hybrid architecture with selective state-space modeling for uncertainty-aware brain tumor segmentation.

PloS one·2026
Same journal

Multicenter evaluation of BACT-Info. and an infection algorithm using Urine Flow Cytometry among clinically diagnosed UTI patients in Indonesia.

PloS one·2026
Same journal

Cross-cultural adaptation and psychometric properties study of Prolonged Grief Disorder Questionnaire (PG-12-R) for caregivers of terminal cancer patients, Thai version.

PloS one·2026
Same journal

Design and in silico validation of donor DNA for RNA-guided recombinase-mediated knockout of mstnb gene in Labeo rohita.

PloS one·2026
Same journal

ViT-MultiRAGNet: A scalable and reliable retrieval-augmented Vision Transformer framework for memory-guided feature fusion multi-modal mammogram classification.

PloS one·2026

相关实验视频

Updated: Jun 9, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

475

DPNet:基于双视角CNN转换器的场景文本检测

Yuan Li1

  • 1School of Physics and Electronic-Electrical Engineering, ABA Teachers University, Wenchuan, Aba Tibetan and Qiang Autonomous Prefecture, Sichuan, China.

PloS one
|October 21, 2024
PubMed
概括

这项研究引入了一种新的双视角CNN转换器模型用于场景文本检测,增强复杂图像的特征提取. 拟议的方法显著提高了多个数据集的检测准确性和稳定性.

科学领域:

  • 计算机视觉 计算机视觉
  • 深度学习 (Deep Learning) 是一种深度学习.
  • 人工智能的人工智能

背景情况:

  • 由于复杂的背景和各种文本外观,场景文本检测具有挑战性.
  • 卷积神经网络 (CNN) 捕捉了局部特征,但缺乏全球背景.
  • 变压器擅长捕获全球图像信息,提供一种互补的方法.

研究的目的:

  • 为改进场景文本检测提出一种新的双视角CNN转换器模型.
  • 增强对文本的全球上下文信息和位置关系的学习.
  • 解决在复杂场景中检测小或复杂文本的挑战.

主要方法:

  • 将道增强自我注意模块 (CESAM) 和空间增强自我注意模块 (SESAM) 集成到ResNet骨干中.
  • 开发一个功能解码器,以完善文本信息和增强细节感知.
  • 使用混合CNN转换器架构进行全面的特征提取.

主要成果:

  • 对各种文本检测场景的模型稳定性进行了显著改进.
  • 与基线相比,总文本的性能增长为2.51%,ICDAR 2015的性能增长为1.87%,MSRA-TD500数据集的性能增长为3.63%.
  • 在文本检测结果中展示了卓越的视觉效果.

更多相关视频

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.8K
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

370

相关实验视频

Last Updated: Jun 9, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

475
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.8K
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

370

结论:

  • 双视角CNN转换器方法有效地结合了本地和全球特征学习,用于场景文本检测.
  • 建议的注意力模块和功能解码器有助于提高准确性和稳定性.
  • 这种方法为复杂的场景文本识别任务提供了有希望的进步.