VTFusion:一个视觉-文本多式融合网络,用于几次拍摄异常检测
IEEE transactions on cybernetics
|January 21, 2026
概括
本研究介绍了VTFusion,这是一个在工业环境中用于少数射击异常检测 (FSAD) 的新框架. 通过适应特定领域任务的视觉和文本特征,VTFusion提高了准确性,提高了工业检查可靠性.
科学领域:
- 计算机视觉 计算机视觉
- 机器学习 机器学习
- 工业自动化 工业自动化
背景情况:
- 短射击异常检测 (FSAD) 需要在有限的正常数据下识别缺陷.
- 当前的方法经常使用一般的图像特征,缺少工业特点.
- 现有的视觉-文本融合策略与语义 misalignment 和交叉模式干扰作斗争.
研究的目的:
- 为工业FSAD开发一个视觉文本多式联络融合框架 (VTFusion).
- 解决当前FSAD方法中的域间隙和语义错位问题.
- 提高工业检查中异常检测的稳定性和准确性.
主要方法:
- 引入了用于视觉和文本数据的自适应特征提取器,以学习特定域的表示.
- 生成合成异常以改善特征可区分性.
- 开发了一种多式预测融合模块,配备了融合块和细分网络,用于像素级异常映射.
主要成果:
- 在MVTec AD (96.8%AUROC) 和Visa (86.2%AUROC) 上,在2次射击的FSAD场景中取得了高性能.
- 在现实世界的工业汽车塑料零部件数据集上,AUPRO达93.5%的实际适用性被证明.
- 在苛刻的工业环境中,VTFusion显著提高了FSAD性能.
结论:
- VTFusion有效地弥合了域差距,并克服了多模式FSAD中的语义 misalignment.
- 拟议的自适应特征提取和融合策略提高了检测准确性和稳定性.
- VTFusion显示出对现实世界工业检查应用的巨大潜力.
相关概念视频
Vision
59.5K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.5K
Nuclear Fusion
33.7K
The process of converting very light nuclei into heavier nuclei is also accompanied by the conversion of mass into large amounts of energy, a process called fusion. The principal source of energy in the sun is a net fusion reaction in which four hydrogen nuclei fuse and ultimately produce one helium nucleus and two positrons.
A helium nucleus has a mass that is 0.7% less than that of four hydrogen nuclei; this lost mass is converted into energy during the fusion. This reaction produces about...
A helium nucleus has a mass that is 0.7% less than that of four hydrogen nuclei; this lost mass is converted into energy during the fusion. This reaction produces about...
33.7K
Color Vision
1.4K
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
1.4K
Protein Networks
4.5K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.5K
Protein Networks
2.8K
2.8K
Network Covalent Solids
16.1K
Network covalent solids contain a three-dimensional network of covalently bonded atoms as found in the crystal structures of nonmetals like diamond, graphite, silicon, and some covalent compounds, such as silicon dioxide (sand) and silicon carbide (carborundum, the abrasive on sandpaper). Many minerals have networks of covalent bonds.
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
16.1K


