Related Experiment Videos
CAPTR-GTP: Class-aware prompting and token refinement with graph token propagation for few-shot ViTs.
Mohammed Al-Habib1, Zuping Zhang1, Abdulrahman Noman1
1School of Computer Science and Engineering, Central South University, Changsha, 410083, Hunan, China.
Summary
Vision Transformers (ViTs) struggle with few-shot learning due to background noise and information loss. CAPTR-GTP, a novel framework, enhances ViTs by refining tokens and propagating class information, achieving state-of-the-art few-shot performance.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Vision Transformers (ViTs) excel at modeling long-range dependencies but face challenges in few-shot learning scenarios.
- Existing ViT adapters often discard or overcompress token information, leading to performance degradation, especially in cluttered scenes or with limited data.
Purpose of the Study:
- To introduce CAPTR-GTP, a novel framework designed to overcome the limitations of Vision Transformers in episodic few-shot learning.
- To preserve and effectively propagate token information, enhancing the model's ability to learn from scarce labeled data.
Main Methods:
- Utilized a pre-trained ViT encoder, largely frozen, for stable patch embeddings.
- Implemented an uncertainty-aware module with Monte Carlo dropout to estimate token reliability and a variance-consistent gate for noisy patches.
- Developed Class-Aware Prompt Refinement to align attention with class semantics and adapt to intra-class variations.
- Introduced a bi-level hierarchical attention module and Graph Token Propagation for refining tokens and aggregating context across clusters.
Main Results:
- CAPTR-GTP demonstrated state-of-the-art or consistently competitive performance across four in-domain and four cross-domain benchmarks.
- The framework achieved strong results under both 5-way 1-shot and 5-shot learning protocols.
- Preservation and propagation of token information proved effective in mitigating information loss and improving few-shot accuracy.
Conclusions:
- CAPTR-GTP offers a robust solution for few-shot learning with Vision Transformers by preserving and refining token information.
- The proposed methods effectively address issues like background noise amplification and reliance on static prototypes.
- The framework shows significant promise for advancing few-shot learning capabilities in computer vision.
Related Concept Videos
Observational Learning
1.3K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.3K
Vision
61.7K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
61.7K
Improving Translational Accuracy
3.8K
3.8K
Improving Translational Accuracy
15.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.6K
Transformers in Distribution System
636
Transformers in distribution systems can be broadly categorized into distribution substation transformers and other distribution transformers. They are crucial for stepping down high transmission voltages to levels suitable for distribution and end-user applications.
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
636
Types Of Transformers
1.8K
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
1.8K