Multimodal Cross-City Semantic Segmentation Based on Similarity-Inspired Fusion and Invertible Transformation