Related Experiment Videos
Lumina-Net: temporal learning with feature fusion for endoscopic artifact removal
Tianjun Yang1, Xingfeng Xu1, Xin Chen2
1Key Laboratory of Mechanism Theory and Equipment Design of Ministry of Education, Tianjin University, Tianjin, 300354, China.
Abstract:
Gastrointestinal (GI) endoscopy is widely used for diagnosis and minimally invasive therapy, yet specular reflections frequently saturate tissue appearance and introduce temporal flicker, degrading both clinical inspection and downstream computational analysis. This study proposes Lumina-Net, a mask-guided video inpainting framework whose spatiotemporal Transformer with overlapping tokens aggregates complementary information across consecutive frames. The decoder incorporates two lightweight modules: Variance-Guided Feature Modulation (VGFM), which recalibrates features using channel statistics to handle mixed-scale specular highlights, and a parameter-free Multi-Scale Energy Free-Space Attention (MS-EFSA) mechanism that derives spatial weights from feature energy to preserve mucosal structures. On the HyperKvasir and GastroHUN datasets, Lumina-Net achieves a peak signal-to-noise ratio (PSNR) of 30.20 dB and reduces the mean squared error (MSE) by approximately 5.3% compared with the strongest baseline, while running at 27 frames per second (FPS) on a single GPU. In a blinded evaluation, clinical experts consistently prefer the visual quality of the proposed method, and improved monocular depth estimation on processed sequences further demonstrates downstream utility. These results indicate that Lumina-Net provides temporally stable specular reflection removal for reliable clinical observation and downstream tasks such as robotic-assisted intervention.