A Multi-domain and Multi-modal Representation Disentangler for Cross-Domain Image Manipulation and Classification