Text guided cross attentive multimodal learning with visual feature modulation for automated skin lesion detection

P Suresh1, P Keerthika2, A R Nitesh Kumar1

  • 1School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, India.

Scientific Reports
|April 6, 2026
PubMed
Summary

Integrating clinical text with dermoscopic images significantly improves automated skin lesion detection. The Text-Guided Cross-Attentive Visual Feature Network (TG-CAVNet) enhances diagnostic accuracy and interpretability in dermatology AI.