Related Experiment Video
Updated: Jun 6, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Reconsidering learnable fine-grained text prompts for few-shot anomaly detection in visual-language models
Delong Han1, Luo Xu1, Mingle Zhou1
1Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences), Jinan, 250353, China; Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, 250014, China.
This study introduces a novel framework for Few-Shot Anomaly Detection (FSAD) in industrial settings, utilizing fine-grained learnable text prompts. The approach enhances accuracy and generalization in identifying industrial defects with limited data.
Area of Science:
- Computer Vision
- Machine Learning
- Industrial Quality Control
Background:
- Few-Shot Anomaly Detection (FSAD) is critical for industrial image analysis with limited training samples.
- Current FSAD methods often rely on extensive manual prompt engineering for visual-language models, limiting accuracy and generalization.
- Domain gaps between training and testing datasets hinder the performance of text prompts in cross-dataset detection.
Purpose of the Study:
- To develop a unified framework for industrial FSAD using fine-grained learnable text prompts.
- To improve the accuracy and robustness of anomaly detection in industrial images.
- To overcome limitations of manual prompt engineering and domain shift in FSAD.
Main Methods:
- Proposed a Fine-grained Text Prompts Adapter (FTPA) with a registration loss to optimize text prompts for capturing fine-grained semantic information.
- Introduced a Dynamic Modulation Mechanism (DMM) to adaptively modulate image and text prompt branches, mitigating post-training errors and domain agnostic issues.
- Integrated FTPA and DMM into a visual-language model framework for enhanced FSAD.
Main Results:
- Achieved state-of-the-art performance in few-shot industrial anomaly detection and segmentation.
- Demonstrated high accuracy with AUROC scores of 98.3% for classification and 96.3% for segmentation in 4-shot settings on MVTec-AD.
- Achieved 93.8% AUROC for classification and 97.9% for segmentation in 4-shot settings on VisA dataset.
Conclusions:
- The proposed fine-grained learnable text prompt framework significantly enhances FSAD performance in industrial applications.
- The FTPA and DMM effectively address challenges of limited data, manual prompt limitations, and domain generalization.
- The method offers a robust and accurate solution for industrial anomaly detection and segmentation tasks.
More Related Videos
Related Concept Videos
Detection of Gross Error: The Q Test
Nonconscious Mimicry
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

