Related Experiment Video
Updated: Jun 1, 2025

07:22
Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
Published on: December 1, 2023
448
Taming vision transformers for clinical laryngoscopy assessment
Xinzhu Zhang1, Jing Zhao1, Daoming Zong1
1School of Computer Science and Technology, East China Normal University, North Zhongshan Road 3663, Shanghai, 200062, China.
Journal of Biomedical Informatics
|January 19, 2025
Summary
MedFormer, a Vision Transformer model, significantly improves early detection of laryngeal cancer (LCA) and precancerous lesions by analyzing laryngoscopic images. It outperforms other models and matches physician evaluations, aiding clinical diagnosis.
Area of Science:
- Otolaryngology
- Artificial Intelligence
- Medical Imaging
Background:
- Laryngoscopy is crucial for diagnosing laryngeal cancer (LCA), but suffers from high inter-observer variability and diagnostic challenges in differentiating precancerous from early cancerous lesions.
- Current diagnostic methods rely heavily on endoscopist expertise, leading to potential inconsistencies and delayed detection.
Purpose of the Study:
- To enhance laryngoscopic image analysis for improved early screening and detection of laryngeal cancer and precancerous conditions.
- To develop a robust AI model that overcomes data scarcity and achieves strong generalization capabilities.
Main Methods:
- Proposing MedFormer, a novel laryngeal cancer classification method utilizing the Vision Transformer (ViT) architecture.
- Implementing a customized transfer learning approach with pre-trained transformers to address data limitations and enable robust out-of-domain generalization.
- Fine-tuning a minimal set of additional parameters for efficient model adaptation.
Main Results:
- MedFormer achieved high sensitivity-specificity values: 98%-89% for precancerous lesions and 89%-97% for cancer detection.
- The model significantly outperformed Convolutional Neural Network (CNN) counterparts and other Vision Transformer (ViT)-based models.
- MedFormer matched or exceeded physician visual evaluations (PVE) in specific scenarios and demonstrated interpretability through visualizations.
Conclusions:
- Visual transformers show significant potential for clinical laryngoscopic assessments.
- MedFormer represents an effective AI-driven method for the early detection of laryngeal cancer, improving diagnostic accuracy and consistency.

