Holistic-Guided Disentangled Learning With Cross-Video Semantics Mining for Concurrent First-Person and Third-Person

Summary

This study introduces a new dataset and method for recognizing concurrent first- and third-person activities (CFT-AR) from wearable cameras. The approach effectively handles temporal and appearance differences, improving environmental understanding for camera wearers.