Parallel Multi-Attention and Gated Fusion for Visual Question Localized Answering in Surgical Scenes