Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

COX-2 inhibition improves immune system homeostasis and decreases liver damage in septic rats.

The Journal of surgical research·2009
Same author

Mass spectral characterization of organophosphate-labeled, tyrosine-containing peptides: characteristic mass fragments and a new binding motif for organophosphates.

Journal of chromatography. B, Analytical technologies in the biomedical and life sciences·2009
Same author

3D-SURFER: software for high-throughput protein surface comparison and analysis.

Bioinformatics (Oxford, England)·2009
Same author

Total arch replacement with stented elephant trunk technique: a proposed treatment for complicated Stanford type B aortic dissection.

Journal of cardiac surgery·2009
Same author

Top-emitting white organic light-emitting devices with a one-dimensional metallic-dielectric photonic crystal anode.

Optics letters·2009
Same author

[Detection of tick and tick-borne pathogen in some ports of Inner Mongolia].

Zhonghua liu xing bing xue za zhi = Zhonghua liuxingbingxue zazhi·2009

Related Experiment Video

Updated: Jun 24, 2025

Author Spotlight: Exploring the Link Between Time Perception of Visual Stimuli and Reading Skills
09:27

Author Spotlight: Exploring the Link Between Time Perception of Visual Stimuli and Reading Skills

Published on: January 19, 2024

1.2K

Towards Visual-Prompt Temporal Answer Grounding in Instructional Video.

Shutao Li, Bin Li, Bin Sun

    IEEE Transactions on Pattern Analysis and Machine Intelligence
    |June 7, 2024
    PubMed
    Summary

    This study introduces a new visual-prompt text span localization method to improve temporal answer grounding in instructional videos (TAGV). The approach enhances locating visual answers by integrating video subtitles and visual prompts, outperforming existing methods.

    More Related Videos

    Using Rapid Serial Visual Presentation to Measure Set-Specific Capture, a Consequence of Distraction While Multitasking
    05:58

    Using Rapid Serial Visual Presentation to Measure Set-Specific Capture, a Consequence of Distraction While Multitasking

    Published on: August 29, 2018

    8.9K
    Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
    09:27

    Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language

    Published on: October 13, 2018

    10.0K

    Related Experiment Videos

    Last Updated: Jun 24, 2025

    Author Spotlight: Exploring the Link Between Time Perception of Visual Stimuli and Reading Skills
    09:27

    Author Spotlight: Exploring the Link Between Time Perception of Visual Stimuli and Reading Skills

    Published on: January 19, 2024

    1.2K
    Using Rapid Serial Visual Presentation to Measure Set-Specific Capture, a Consequence of Distraction While Multitasking
    05:58

    Using Rapid Serial Visual Presentation to Measure Set-Specific Capture, a Consequence of Distraction While Multitasking

    Published on: August 29, 2018

    8.9K
    Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
    09:27

    Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language

    Published on: October 13, 2018

    10.0K

    Area of Science:

    • Computer Science
    • Artificial Intelligence
    • Video Analysis

    Background:

    • Temporal answer grounding in instructional video (TAGV) is crucial for video understanding.
    • Existing methods struggle with weak correlations between text questions and visual answers in TAGV.
    • Current visual span-based predictors are insufficient for accurate TAGV.

    Purpose of the Study:

    • To propose a novel Visual-Prompt Text Span Localization (VPTSL) method for TAGV.
    • To improve the semantic understanding between textual queries and visual answers in instructional videos.
    • To reformulate TAGV as a subtitle span localization task enhanced by visual prompts.

    Main Methods:

    • Developed a VPTSL method incorporating timestamped subtitles and learnable visual prompt embeddings.
    • Utilized a pre-trained language model to learn joint semantic representations from text, subtitles, and visual prompts.
    • Reformulated TAGV as locating subtitle spans corresponding to visual answers.

    Main Results:

    • The VPTSL method significantly outperforms state-of-the-art methods on five benchmark datasets (MedVidQA, TutorialVQA, VehicleVQA, CrossTalk, Coin).
    • Achieved substantial improvements in mean Intersection over Union (mIoU) scores.
    • Demonstrated the effectiveness of the visual prompt and text span-based predictor integration.

    Conclusions:

    • The proposed VPTSL method offers a more effective approach to temporal answer grounding in instructional videos.
    • Integrating visual prompts and subtitle information enhances the model's ability to ground answers visually.
    • This work advances the field of instructional video analysis and question answering.