Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Observational Learning01:12

Observational Learning

210
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
210
Force Classification01:22

Force Classification

1.3K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.3K
Introduction to Learning01:18

Introduction to Learning

472
Learning is the process of acquiring knowledge or skills through practice or experience, leading to long-lasting behavioral changes. This acquisition occurs through interaction with the environment and requires practice or experience. For instance, mastering a skill such as surfing requires considerable practice and experience, highlighting the essential role of repeated interactions with the environment in learning.
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
472
Vision01:24

Vision

53.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Synergistic effects of organic salt curing and heating methods on quality and flavor preservation of large yellow croaker during frozen storage.

Food chemistry·2026
Same author

Identification of diagnostic markers for diabetic kidney disease by weighted gene co‑expression network analysis and machine learning.

International journal of molecular medicine·2026
Same author

Linoleic Acid Potentiates Response to Chemotherapy in Biliary Tract Cancer Through RARγ Activation.

FASEB journal : official publication of the Federation of American Societies for Experimental Biology·2025
Same author

New power in cancer immunotherapy: the rise of chimeric antigen receptor macrophage (CAR-M).

Journal of translational medicine·2025
Same author

Circular RNA-Based Molecular Computation Enhances Plasma Biomarker Detection in Biliary Tract Cancer.

Angewandte Chemie (International ed. in English)·2025
Same author

Molecular epidemiological characteristics, variant spectrum and genotype-phenotype correlation of glucose-6-phosphate dehydrogenase deficiency in China: A population-based multicenter study using newborn screening.

PloS one·2024

Related Experiment Video

Updated: Jul 21, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

572

Fine-Grained Visual-Text Prompt-Driven Self-Training for Open-Vocabulary Object Detection.

Yanxin Long, Jianhua Han, Runhui Huang

    IEEE Transactions on Neural Networks and Learning Systems
    |July 28, 2023
    PubMed
    Summary

    This study introduces a novel fine-grained visual-text prompt-driven self-training method for open-vocabulary object detection. The approach enhances vision-language models (VLMs) for better instance localization and achieves state-of-the-art results on unseen classes.

    More Related Videos

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K
    A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
    04:23

    A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

    Published on: April 21, 2023

    1.9K

    Related Experiment Videos

    Last Updated: Jul 21, 2025

    Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
    03:31

    Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

    Published on: December 15, 2023

    572
    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K
    A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
    04:23

    A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

    Published on: April 21, 2023

    1.9K

    Area of Science:

    • Computer Vision
    • Artificial Intelligence
    • Machine Learning

    Background:

    • Vision-language models (VLMs) show promise in zero-shot classification.
    • Extending VLMs to object detection often uses self-training with pseudolabels for unseen classes.
    • Current VLMs lack fine-grained alignment crucial for object instance detection due to global image-text pretraining.

    Purpose of the Study:

    • To propose a fine-grained visual-text prompt-driven self-training paradigm for open-vocabulary object detection (VTP-OVD).
    • To enhance the fine-grained alignment capabilities of VLMs for object detection tasks.
    • To improve the performance of detecting objects from unseen classes.

    Main Methods:

    • Introduced a fine-grained visual-text prompt adapting stage to improve VLM alignment.
    • Employed learnable text prompts for an auxiliary dense pixelwise prediction task to achieve fine-grained alignment.
    • Developed a visual prompt module to supply prior task information to the vision branch for better adaptation.

    Main Results:

    • The proposed VTP-OVD method significantly enhances fine-grained alignment in VLMs.
    • Achieved state-of-the-art performance in open-vocabulary object detection.
    • Demonstrated strong results, including 31.5% mAP on unseen classes of the COCO dataset.

    Conclusions:

    • The fine-grained prompt adaptation is effective for improving VLMs in open-vocabulary object detection.
    • The VTP-OVD paradigm offers a powerful approach for detecting objects from unseen classes.
    • This work advances the capabilities of VLMs for complex downstream vision tasks.