A Vision-Language-Guided Multimodal Fusion Network for Glottic Carcinoma Early Diagnosis: Model Development and

Zhaohui Jin1, Yi Shuai2, Yun Li2

  • 1College of Big Data and Internet, Shenzhen Technology University, Pingshan District, 3002 Lantian Road, Shenzhen, Guangdong, 518118, China, 86 19276679344.

JMIR Medical Informatics
|October 8, 2025
PubMed
Summary

A new vision-language model, VLMF-Net, improves early diagnosis of glottic carcinoma (GC) by integrating text and images. This AI tool shows superior accuracy and robustness compared to existing methods, aiding early detection efforts.