Related Experiment Video
Updated: Jul 3, 2026

06:08
Quantitative Visualization and Detection of Skin Cancer Using Dynamic Thermal Imaging
Published on: May 5, 2011
17.3K
Text guided cross attentive multimodal learning with visual feature modulation for automated skin lesion detection
P Suresh1, P Keerthika2, A R Nitesh Kumar1
1School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, India.
Scientific Reports
|April 6, 2026
Summary
Integrating clinical text with dermoscopic images significantly improves automated skin lesion detection. The Text-Guided Cross-Attentive Visual Feature Network (TG-CAVNet) enhances diagnostic accuracy and interpretability in dermatology AI.
Area of Science:
- Artificial Intelligence
- Dermatology
- Medical Imaging
Background:
- Automated skin lesion detection is crucial for early diagnosis but current deep learning models often overlook clinical context, limiting robustness and interpretability.
- Visually ambiguous skin lesions pose a significant challenge for image-only diagnostic systems.
Purpose of the Study:
- To develop an explainable multimodal framework integrating clinical text and dermoscopic images for improved automated skin lesion detection.
- To enhance diagnostic accuracy and semantic grounding in AI-powered dermatological analysis.
Main Methods:
- Proposed a Text-Guided Cross-Attentive Visual Feature Network (TG-CAVNet) architecture combining Bio-ClinicalBERT for text encoding and EfficientNet-B4 for visual feature extraction.
- Employed text-guided channel-wise feature modulation, text-queried cross-attention for semantic-spatial alignment, and adaptive multi-stream fusion.
- Trained the model end-to-end using hybrid focal and cross-entropy losses on a dataset of 6194 aligned image-text samples.
Main Results:
- TG-CAVNet achieved 90.75% accuracy and a macro Jaccard score of 0.82, outperforming state-of-the-art multimodal baselines.
- Ablation studies confirmed the significant contributions of individual components and their synergistic effects.
- Attention visualizations demonstrated improved model interpretability.
Conclusions:
- Text-guided cross-attentive multimodal learning enhances both performance and explainability in automated skin lesion identification.
- The TG-CAVNet framework highlights the importance of integrating clinical context with visual data for robust dermatological AI decision support.
