Related Experiment Videos
Real-time target recognition with a hybrid multi-attention transformer-CNN (AttnConvNeXt) for computational ghost
None:
A hybrid multi-head attention transformer-CNN (AttnConvNeXt) model for computational ghost imaging (CGI) that can recognize targets both with and without images is presented in this paper. This unified architecture runs across several resolutions (128×128, 64×64, and 32×32) and directly analyzes raw 1D bucket measurements without reconstruction, in contrast to previous GI classifiers that were restricted to either reconstructed images or pre-processed signals. AttnConvNeXt achieves robust classification under low sampling ratios (SR=0.8) where traditional approaches fail by combining multi-head attention with convolutional layers to capture both local features and global dependencies. Our model achieves 99%-100% recognition over resolutions when used for reconstructed images, providing a high-accuracy baseline. It outperforms a 12-layer CNN by 65% when processing solely on bucket signals in image-free mode achieving 84% accuracy at SR<1. Real-time viability is demonstrated by the recognition time scaling effectively with resolution from 0.017 s/image (128×128) to 0.00056 s/signal (raw measurements). By creating the first multi-resolution, dual-mode GI recognition framework, to the best of our knowledge, this study removes the need for Fourier transforms for reconstruction-based recognition and makes it possible to use it for medical diagnostics, low-light surveillance, and scattering media.