Related Experiment Video
Updated: May 6, 2026

Three-dimensional Optical-resolution Photoacoustic Microscopy
Published on: May 3, 2011
Sound-field-projection synthesis using latent diffusion model for acousto-optic reconstruction
Risako Tanigawa1,2, Kenji Ishikawa1, Noboru Harada1
1Communication Science Laboratories, NTT, Inc., 3-1 Morinosato-Wakamiya, Atsugi, Kanagawa 243-0198, Japan.
Abstract:
Acousto-optic sensing (AOS) is a non-contact method for measuring sound using light, suitable for environments where microphones are impractical, such as confined spaces or within airflow. Despite its effectiveness, AOS captures line-integrated sound pressure along the optical path, resulting in signals that cannot be interpreted as those from point-wise microphones. To obtain the sound pressure distribution in three-dimensional (3D) space, volumetric sound-field reconstruction is required. Existing methods require multi-directional line-integrated projection data, typically obtained either by deploying multiple devices or by repeatedly exciting the sound source. The former requires deploying multiple high-cost devices, which is often impractical, whereas the latter is not applicable to sound sources that cannot be reproduced consistently. To overcome these limitations, we propose a task called sound projection synthesis, which synthesizes data in specified directions based on observation data. We achieve this by using a latent diffusion model conditioned on observed projections and view angles to generate new sound-field projections. A model pretrained on natural images was fine-tuned using sound-field data and optimized with a pixel-wise loss. Experiments demonstrate that the model can generate realistic projection data. Combining nine observed views with nine generated views improved 3D reconstruction accuracy compared with using only nine observed views.

