Related Experiment Video
Updated: Jul 1, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.4K
A survey on advancements in image-text multimodal models: From general techniques to biomedical implementations
Ruifeng Guo1, Jingxuan Wei1, Linzhuang Sun1
1Shenyang Institute of Computing Technology, Chinese Academy of Sciences, Shenyang, 110168, China; University of Chinese Academy of Sciences, Beijing, 100049, China.
Computers in Biology and Medicine
|June 15, 2024
Summary
This survey reviews image-text multimodal models, detailing their technological evolution and impact on biomedical applications. It analyzes general model advancements and domain-specific challenges, offering solutions for future research.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Computer Vision
Background:
- Large Language Models (LLMs) have spurred advancements in Natural Language Processing (NLP).
- Image-text multimodal models integrate visual and textual data, gaining significant research attention.
- Existing surveys often overlook the influence of general models on domain-specific advancements.
Purpose of the Study:
- To review the technological evolution of image-text multimodal models.
- To analyze the impact of general model development on biomedical multimodal technologies.
- To identify and address challenges in general and domain-specific multimodal model development.
Main Methods:
- Reviewing technological evolution from feature spaces to large model architectures.
- Analyzing common components, tasks, and challenges of image-text multimodal models.
- Summarizing general model architectures, components, and data, with a focus on biomedical applications.
Main Results:
- General image-text multimodal technologies significantly promote biomedical multimodal progress.
- Biomedical datasets present unique importance and complexity for model development.
- Challenges are categorized into 2 external and 5 intrinsic factors, with proposed solutions.
Conclusions:
- Understanding the evolution of general models is crucial for domain-specific research.
- Targeted solutions are provided for challenges in multimodal model development and application.
- This survey guides future research directions in image-text multimodal AI.

