Improving Medical Visual Representation Learning With Pathological-Level Cross-Modal Alignment and Correlation