PRORED: a hybrid transformer framework with progressive refinement decoding for segmenting dynamic speech MRI
Ying He1,2, Qianni Zhang1,2, Marc E Miquel3,4
1School of Electronic Engineering and Computer Science, Queen Mary University of London, London E1 4NS, United Kingdom.
Objectives:
Dynamic MRI of the upper vocal tract is increasingly used to study speech. Image segmentation is often required to analyse the organs of speech; however, manual segmentation is labour intensive and time consuming and automatic methods are being developed. In this paper, a new hybrid transformer network is proposed for such task.
Methods:
We introduce a deep learning-based decoder model termed "Progressively Refinement Decoding (PRORED)." This model incorporates a directional field (DF) module designed to capture the contour details of features. The acquired contour information is leveraged to refine the boundaries both between and within classes. By integrating the DF module at different stages of the decoder, features are enhanced progressively, ensuring a more detailed and accurate segmentation.
Results:
Our model is evaluated using a publicly accessible speech MRI dataset and a cardiac dataset. The metrics employed are the Dice coefficient and the Hausdorff distance. Results indicate that our model attains an average Dice coefficient of 97.78 and a Hausdorff distance of 6.84 mm. Additionally, our network was able to identify closure patterns more efficiently than the baseline network and previously published work. In addition, the model was also evaluated on a cardiac dataset, and achieved 91.90 dice score.
Conclusions:
The proposed model leads to a more accurate segmentation of speech MRI data and in particular allows for a better velopharyngeal closure study. The proposed model was also evaluated on a cardiac dataset and achieved competitive performance, showing its strong generalizability.
Advances In Knowledge:
First model that utilizes vision transformer and progressive refinement decoder to segment dynamic speech MRI.


