场景图和基于自然语言的语义图像检索使用视觉传感器数据
1Department of Computer Engineering, Keimyung University, Daegu 42601, Republic of Korea.
Sensors (Basel, Switzerland)
|September 19, 2025
概括
本研究引入了一种新的图形神经网络 (GNN) 方法,用于基于文本的图像检索,通过比较语义和场景图来提高准确性,以便更好地理解视觉内容.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 基于文本的图像检索通常使用关键字匹配,这与语义细微差别和有限的查询信息有关.
- 当前的方法在处理复杂的场景或新的句子时缺乏准确性,无法捕捉完整的上下文含义.
研究的目的:
- 开发一种基于文本的图像检索的新方法,克服关键字匹配的局限性.
- 通过使文字描述和视觉场景内容之间的定量比较来提高检索准确度.
主要方法:
- 将句子转化为语义图表,将图像转化为场景图表.
- 使用图形神经网络 (GNN) 来学习节点和边缘特征,生成图形嵌入用于比较.
- 实施一个对比的GNN框架,使用硬负挖矿来匹配语义和场景图.
主要成果:
- 提出的基于GNN的方法在Visual Genome数据集上获得了0.745的nDCG@50最高得分.
- 与随机采样和全图表相比,显示了大约7.7个百分点的改善.
- 通过结构性地解释复杂的场景,成功地获取了语义上相关的图像.
结论:
- 这种基于GNN的新方法有效地解决了基于文本的图像检索的局限性.
- 语义和场景图的定量比较显著提高了检索准确度.
- 通过图形嵌入进行场景的结构解释,可以从自然语言查询中进行强大的图像检索.
相关概念视频
Visual System
1.7K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.7K
Vision
59.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.4K


