A survey on advancements in image-text multimodal models: From general techniques to biomedical implementations

Ruifeng Guo1, Jingxuan Wei1, Linzhuang Sun1

  • 1Shenyang Institute of Computing Technology, Chinese Academy of Sciences, Shenyang, 110168, China; University of Chinese Academy of Sciences, Beijing, 100049, China.

Summary

This survey reviews image-text multimodal models, detailing their technological evolution and impact on biomedical applications. It analyzes general model advancements and domain-specific challenges, offering solutions for future research.