Multi-Modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training

Related Concept Videos