Related Experiment Video
Updated: Jan 15, 2026

04:25
Author Spotlight: Bridging Gaps in Anatomy and Establishing a Foundation for Algorithmic Studies
Published on: December 15, 2023
3.7K
Text-guided multimodal deep learning in magnetic resonance imaging for spinal structures segmentation and lumbar
Jingxin Liu1,2, Xinran Zhu1,2, Zhangzhen Shi3
1Department of Radiology, China-Japan Union Hospital of Jilin University, Changchun, China.
Quantitative Imaging in Medicine and Surgery
|October 13, 2025
Summary
This study introduces a text-guided deep learning method for lumbar MRI analysis, improving spinal structure segmentation and abnormality identification. The approach enhances diagnostic accuracy and efficiency for radiologists.
Area of Science:
- Medical Imaging Analysis
- Deep Learning
- Radiology
Background:
- Unimodal deep learning (DL) methods struggle to model relationships between textual semantics and medical images.
- Lumbar MRI analysis requires effective modeling of latent relationships within textual descriptions for accurate target identification.
- Existing methods are insufficient for comprehensive analysis of spinal structures and abnormalities in lumbar MRI scans.
Purpose of the Study:
- To propose a text-guided multimodal DL method for segmenting spinal structures and identifying lumbar abnormalities in T1WI and T2WI MRI scans.
- To leverage textual semantics and cross-modal features for improved accuracy in lumbar MRI analysis.
- To develop an automated workflow that assists radiologists and reduces workload.
Main Methods:
- Employed ConvNeXt V2 as the image encoder and a Contrastive Language-Image Pretraining (CLIP)-based text encoder.
- Utilized self-supervised pretraining for each encoder on unlabeled lumbar MRI scans and clinical reports.
- Developed a text-guided DL method integrating both encoders on 201 fully annotated T1WI and T2WI MRI samples.
Main Results:
- Achieved a mean Intersection over Union (mIoU) of 0.823±0.053 for segmenting 19 spinal structures, outperforming existing methods.
- Demonstrated high performance in lumbar abnormality identification with a recall of 0.867±0.027, low FPR of 0.079±0.015, and precision of 0.893±0.028.
- Showcased no significant difference in segmentation metrics between upper and lower Dice scores (P=0.744).
Conclusions:
- The proposed text-guided DL method effectively segments spinal structures and identifies lumbar abnormalities in multi-sequence MRI scans.
- Integration of cross-modal features significantly improved segmentation accuracy and abnormality identification performance.
- The method shows clinical potential for automated workflows, aiding diagnosis and reducing radiologist workload.
