A Dual-Attention Learning Network With Word and Sentence Embedding for Medical Visual Question Answering