Related Experiment Video
Updated: Aug 5, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Language-Guided Segmentation of Medical Images: A Review of Foundation Models
Saqib Qamar1,2
1Department of Intelligent Systems, KTH Royal Institute of Technology, 10044 Stockholm, Sweden.
Bioengineering (Basel, Switzerland)
|July 28, 2026
Summary
Vision-language foundation models revolutionize medical image segmentation using text prompts for diverse tasks. This survey explores their technical background, taxonomy, adaptation, and clinical applications, highlighting challenges and future directions.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Computer Vision
Background:
- Vision-language foundation models (VLMs) integrate image encoders and text prompts for versatile medical image segmentation.
- These models enable segmentation of various anatomical structures, lesions, and modalities via natural language.
- Recent advancements include contrastive pretraining and models like Segment Anything Model (SAM) and its medical adaptations.
Purpose of the Study:
- To provide a comprehensive survey of VLMs for medical image segmentation.
- To categorize existing VLMs and discuss their adaptation strategies.
- To identify current challenges and propose a future research roadmap.
Main Methods:
- Literature review and synthesis of VLMs in medical image segmentation.
- Categorization of models into text-prompt-guided, LLM-embedded, and hybrid frameworks.
- Analysis of adaptation techniques (fine-tuning, LoRA, adapters, prompt engineering) across different imaging modalities.
Main Results:
- Organized literature by modality (CT, MRI, pathology, radiography, ultrasound) and clinical use (organ segmentation, tumor delineation, radiotherapy planning).
- Summarized evaluation metrics and benchmark datasets.
- Identified key challenges: prompt dependence, mask hallucination, slow inference, and data limitations.
Conclusions:
- VLMs offer significant potential for medical image segmentation, transforming clinical workflows.
- Addressing challenges like prompt dependence and inference speed is crucial for clinical translation.
- Future research should focus on trustworthy deployment, multimodal pretraining, and seamless clinical integration.
