Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

simDP: Sim-to-Real Transfer with Shared Action Spaces.

Sensors (Basel, Switzerland)·2026
Same author

Differential Modulation of Postprandial Glycemic, Incretin, and Satiety Responses by Low-Digestible Carbohydrates in Humans: An Exploratory Investigation.

Nutrients·2026
Same author

Sensing the Action: Rethinking Sensor Modalities and Multi-Modal Fusion in Vision-Language-Action Models for Robotic Manipulation.

Sensors (Basel, Switzerland)·2026
Same author

Ethylene-driven enhancement of bioactive metabolites and in vitro functionality in soybean (Glycine max (L.) Merr.) and mung bean (Vigna radiata (L.) Wilczek) leaves grown in vertical farms: a comparative study.

BMC plant biology·2026
Same author

Enhancement of isoflavone aglycones, GABA, and mineral bioavailability in <i>Apios americana</i> Medikus by co-fermentation with <i>Lactiplantibacillus plantarum</i> LAB02 and <i>Levilactobacillus brevis</i> BMK484.

Food chemistry: X·2026
Same author

Threshold Switching Behavior and Underlying Mechanisms in Pure SiO<sub>2</sub>-Based Selectors.

ACS applied materials & interfaces·2026

相关实验视频

Updated: Jan 13, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

3.3K

通过语言和视觉传感器数据的交叉注意调整进行关键字条件的图像细分.

Hye Rim Kim1, Byoung Chul Ko1

  • 1Department of Computer Engineering, Keimyung University, Daegu 42601, Republic of Korea.

Sensors (Basel, Switzerland)
|October 29, 2025
PubMed
概括

本研究介绍了KeySeg,这是一种基于关键字的图像细分模型. KeySeg通过明确地将语言理解与视觉执行联系起来来改善基于推理的细分,以获得更准确的结果.

科学领域:

  • 计算机科学 计算机科学
  • 人工智能的人工智能
  • 机器学习 机器学习

背景情况:

  • 多模式大语言模型 (LLM) 允许用于图像分割的联合视觉和语言处理.
  • 由于语言解释和细分执行之间的断开,现有的方法面临语义上的差异.

研究的目的:

  • 提出KeySeg,一个新的架构,解决多式模式图像细分中的语义差距.
  • 显式编码并将推断的查询条件集成到细分过程中.

主要方法:

  • KeySeg将多式联网输入的核心概念嵌入到一个[KEY]令牌中.
  • 一个交叉注意力融合模块将[KEY]令牌与[SEG]令牌集成.
  • 关键字对齐损失确保[KEY]令牌与查询的语义核心对齐.

主要成果:

  • KeySeg在细分标准中明确而准确地反映了查询条件.
  • 该模型在条件解释方面表现出更高的准确性.
  • 达到表达能力和解释稳定性,即使在复杂的语言条件下.

结论:

  • 在图像分割中,KeySeg有效地弥合了语言和视觉之间的语义差距.
关键词:
关键词条件 关键词条件 关键词条件以关键字为条件的图像细分 图像细分多模式学习是多模式学习.推理细分分类的推理.视觉传感器 视觉传感器视觉语言模型

更多相关视频

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.6K
Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
07:11

Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping

Published on: December 8, 2023

2.3K

相关实验视频

Last Updated: Jan 13, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

3.3K
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.6K
Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
07:11

Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping

Published on: December 8, 2023

2.3K
  • 分离条件推理和细分指令提高了模型性能.
  • 该架构为基于推理的图像细分提供了稳定而准确的方法.