RefZVC: Refinable Zero-Shot Video Captioning by Test-Time Reinforcement Polishing

Summary

This study introduces Refinable Zero-shot Video Captioning (RefZVC), a novel framework for generating video descriptions without paired video-text data. RefZVC uses test-time reinforcement polishing to refine captions, significantly improving performance on benchmarks.