Related Experiment Video
Updated: Jul 13, 2025

Author Spotlight: Noninvasive Cerebral Blood Flow Determination in Human Functional Brain Region for Diagnosis of Neurological Disorders
Published on: May 31, 2024
Transformer-based deep learning denoising of single and multi-delay 3D arterial spin labeling
Qinyang Shou1, Chenyang Zhao1, Xingfeng Shao1
1Laboratory of Functional MRI Technology (LOFT), Stevens Neuro Imaging and Informatics Institute, University of Southern California, Los Angeles, California, USA.
Purpose:
To present a Swin Transformer-based deep learning (DL) model (SwinIR) for denoising single-delay and multi-delay 3D arterial spin labeling (ASL) and compare its performance with convolutional neural network (CNN) and other Transformer-based methods.
Methods:
SwinIR and CNN-based spatial denoising models were developed for single-delay ASL. The models were trained on 66 subjects (119 scans) and tested on 39 subjects (44 scans) from three different vendors. Spatiotemporal denoising models were developed using another dataset (6 subjects, 10 scans) of multi-delay ASL. A range of input conditions was tested for denoising single and multi-delay ASL, respectively. The performance was evaluated using similarity metrics, spatial SNR and quantification accuracy of cerebral blood flow (CBF), and arterial transit time (ATT).
Results:
SwinIR outperformed CNN and other Transformer-based networks, whereas pseudo-3D models performed better than 2D models for denoising single-delay ASL. The similarity metrics and image quality (SNR) improved with more slices in pseudo-3D models and further improved when using M0 as input, but introduced greater biases for CBF quantification. Pseudo-3D models with three slices achieved optimal balance between SNR and accuracy, which can be generalized to different vendors. For multi-delay ASL, spatiotemporal denoising models had better performance than spatial-only models with reduced biases in fitted CBF and ATT maps.
Conclusions:
SwinIR provided better performance than CNN and other Transformer-based methods for denoising both single and multi-delay 3D ASL data. The proposed model offers flexibility to improve image quality and/or reduce scan time for 3D ASL to facilitate its clinical use.

