Viewpoint-Adaptive Representation Disentanglement Network for Change Captioning

Summary

This study introduces a viewpoint-adaptive network to accurately describe image changes by distinguishing real alterations from pseudo changes caused by viewpoint shifts. The method enhances change captioning performance on multiple datasets.