Related Experiment Videos
Training-free OCT denoising via anatomically guided latent manipulation of a pretrained generative foundation model
Sajed Rakhshani1, Hossein Rabbani1
1Medical Image and Signal Processing Research Center, School of Advanced Technologies in Medicine, Isfahan University of Medical Sciences, Isfahan, Iran.
Abstract:
Optical coherence tomography (OCT) is a cornerstone of retinal diagnostics; however, its clinical utility is often limited by speckle noise that obscures fine anatomical structures. Conventional denoising approaches either rely on handcrafted priors that tend to oversmooth retinal textures or employ supervised deep networks that require paired clean data and may introduce hallucinated features when applied outside their training distribution. In this work, we propose a training-free, anatomically guided OCT denoising framework that leverages a frozen variational autoencoder (VAE) foundation model as a fixed generative prior. The method first identifies speckle-dominant background regions using a Gaussianization transform based on the Bessel-K statistical model, followed by Gaussian mixture modeling to separate retinal and non-retinal structures. The latent encoding of the extracted noise-dominant region provides an estimate of the noise-related direction within the VAE latent space. By using the VAE posterior, a single deterministic encode-directional-suppression-decode operation then shifts the noisy embedding away from this estimated direction, attenuating speckle while preserving retinal layer integrity. Extensive evaluation across six publicly available OCT datasets (SS-OCT, 3D-OCT-1000, Bioptigen, A2A, PKU, and HCIRAN) demonstrates consistent improvements over both model-driven and learning-based baselines. The proposed method achieves a mean PSNR gain of +5.6 dB and an SSIM improvement of +0.32 relative to the strongest supervised comparator, without any retraining, fine-tuning, or iterative optimization. Furthermore, the framework is inference-deterministic, requires only a single forward pass ensuring stability across varying noise levels. These results indicate that controlled in-distribution latent manipulation of pretrained foundation models provides a scalable, computationally efficient, and anatomically faithful alternative to conventional OCT denoising strategies.