Related Experiment Video
Updated: May 29, 2026

Combining Reflectance Confocal Microscopy with Optical Coherence Tomography for Noninvasive Diagnosis of Skin Cancers via Image Acquisition
Published on: August 18, 2022
Multimodal Deep Learning for Benign vs Malignant Eyelid Lesion Classification Using Optical Coherence Tomography and
Weronika Jakubowska1,2, Clément Playout1,2, Renaud Duval1,2
1Department of Ophthalmology, Université de Montréal, Montreal, Quebec, Canada.
Purpose:
To evaluate the performance of deep learning models using optical coherence tomography (OCT) volumes, clinical photographs, and their multimodal fusion to classify eyelid lesions as benign or malignant, using histopathology as the reference standard.
Methods:
Prospective cohort study conducted between January 2023 and January 2025 at a single tertiary academic oculofacial plastic surgery center. A total of 65 patients with 71 periocular lesions undergoing routine biopsy were imaged with spectral-domain OCT and slit-lamp photography before biopsy. Images were processed into 3 Vision Transformer architectures: an OCT model, a photograph model, and a multimodal fusion model integrating both modalities. Classification performance for benign versus malignant lesions was evaluated with 5-fold cross-validation. Performance metrics included sensitivity, specificity, accuracy, kappa, and area under the receiver operating characteristic curve.
Results:
Of 71 lesions, 52% were benign and 48% malignant. The OCT model achieved 73.9% accuracy, 74.4% sensitivity, and an area under the receiver operating characteristic curve of 82.5%. The photograph model reached 82.3% accuracy, 89.6% sensitivity, and an area under the receiver operating characteristic curve of 91.7%. The multimodal fusion model performed best, with 83.2% accuracy, 91.8% sensitivity, 81.1% precision, kappa of 0.68, and an area under the receiver operating characteristic curve of 92.1%.
Conclusions:
This study demonstrates the feasibility of analyzing OCT volumes using deep learning, with diagnostic accuracy improved through multimodal integration with clinical photography. These results suggest that multimodal artificial intelligence may serve as a scalable, noninvasive, and physician-independent tool to distinguish benign from malignant periocular lesions and streamline patient triage. As models evolve to identify histopathologic features more reliably, they may also reduce dependence on incisional biopsy prior to definitive treatment.