Related Experiment Video
Updated: Jun 13, 2026

Optical Frequency Domain Imaging of Ex vivo Pulmonary Resection Specimens: Obtaining One to One Image to Histopathology Correlation
Published on: January 22, 2013
Improving Laryngoscopy Image Analysis Through Integration of Global Information and Local Features in VoFoCD Dataset
Thao Thi Phuong Dao1,2,3,4, Tuan-Luc Huynh1,3, Minh-Khoi Pham5
1University of Science, Ho Chi Minh City, Vietnam.
This study introduces a new dataset and a multitask AI model (MEAL) for analyzing laryngoscopy images. The model accurately detects vocal fold lesions and classifies images, aiding in diagnosing vocal fold disorders.
Area of Science:
- Medical Imaging
- Artificial Intelligence in Healthcare
- Otolaryngology
Background:
- Laryngoscopy is crucial for diagnosing vocal fold disorders, requiring precise identification of anatomical structures and lesions.
- Current diagnostic methods lack simultaneous optimization for object detection and image classification in laryngoscopy.
- Accurate visual assessment is vital for effective clinical decision-making in laryngology.
Purpose of the Study:
- To introduce the VoFoCD dataset for object detection and image classification in laryngoscopy.
- To propose a novel Multitask Efficient trAnsformer network for Laryngoscopy (MEAL) for enhanced diagnosis.
- To improve the accuracy and interpretability of AI-assisted vocal fold disorder diagnosis.
Main Methods:
- Development of the VoFoCD dataset comprising 1724 laryngology images across four classes and six glottic object types.
- Implementation of the MEAL network, a multitask transformer model for classifying vocal fold images and detecting glottic landmarks/lesions.
- Integration of attention maps within MEAL for explainable AI, visualizing critical regions for clinical interpretation.
Main Results:
- The MEAL model achieved high performance on the VoFoCD dataset, with 0.951 accuracy for image classification.
- Object detection performance was strong, yielding a mean average precision of 0.874 (mAP50).
- The model demonstrated robustness in simulated clinical scenarios, including those with laryngoscopy process variations.
Conclusions:
- The proposed MEAL network effectively integrates global and local features for accurate vocal fold image analysis.
- MEAL aids in visually identifying and classifying benign and malignant vocal fold lesions, supporting laryngologists.
- This AI-driven approach enhances diagnostic capabilities in laryngeal endoscopy, improving patient care for vocal fold disorders.
More Related Videos
12:22Multimodal Volumetric Retinal Imaging by Oblique Scanning Laser Ophthalmoscopy oSLO and Optical Coherence Tomography OCT
Published on: August 4, 2018
05:56Author Spotlight: Enhancing Diagnostic Strategies and Biomarker Development for Comprehensive Lung Function Analysis
Published on: August 9, 2024
Related Concept Videos
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids, corniculates, and...
Endoscopic Studies I: Bronchoscopy and Thoracoscopy
Bronchoscopy
Description
Bronchoscopy is a procedure that involves direct visualization of the larynx, trachea, and bronchi for diagnostic and therapeutic purposes. A flexible fiber optic or rigid bronchoscope is used to carry out the procedure. The fiber-optic bronchoscope is more frequently used due to...