Related Experiment Video
Updated: Aug 5, 2026

11:19
Multimodal Hierarchical Imaging of Serial Sections for Finding Specific Cellular Targets within Large Volumes
Published on: March 20, 2018
LIGHT: Learning Image-text Grounding for Hierarchical Tumor Localization and Subtype Classification in Multimodal
IEEE Transactions on Medical Imaging
|July 31, 2026
Summary
This study introduces LIGHT, a novel framework for brain tumor analysis using MRI scans. LIGHT accurately localizes tumors and classifies subtypes by aligning medical images with radiology reports, improving diagnostic precision.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computational Biology
Background:
- Accurate brain tumor analysis requires both precise localization and subtype classification.
- Current vision-only methods often treat localization and classification separately, limiting alignment with clinical reporting semantics.
- Existing approaches struggle to integrate fine-grained anatomical information crucial for radiological interpretation.
Purpose of the Study:
- To develop a 3D vision-language framework, LIGHT, for hierarchical brain tumor analysis.
- To formulate brain tumor localization and subtype classification as an image-text retrieval task within a shared embedding space.
- To enable a unified coarse-to-fine analysis of tumor location and type directly from multimodal MRI data.
Main Methods:
- LIGHT utilizes a hierarchical retrieval vocabulary with 21 coarse regions, 563 fine subregions, and 5 tumor subtypes.
- The framework employs large-scale foundation pretraining on 99,813 MRI-report pairs.
- Grounded task fine-tuning incorporates segmentation-guided tumor crops and multi-template prompt supervision for enhanced alignment.
Main Results:
- LIGHT achieved 72.1% accuracy for coarse-grained localization and 70.1% top-1 accuracy for fine-grained localization across 11,034 cases.
- Subtype classification accuracy reached 76.0%, 83.7%, and 84.6% on three distinct 3D datasets.
- Cross-setting subtype recognition was further validated through additional 2D experiments.
Conclusions:
- Report-grounded image-text retrieval offers a powerful approach for brain tumor MRI analysis.
- LIGHT provides interpretable anatomical and diagnostic outputs, enhancing clinical utility.
- The framework demonstrates the potential of vision-language models in advancing medical image analysis.