Related Experiment Video
Updated: Apr 26, 2026

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
Data Augmentation-Enhanced Myocardial Infarction Classification and Localization Using a ResNet-Transformer Cascaded
Yunfan Chen1, Qi Gao1, Jinxing Ye1
1Hubei Key Laboratory for High-Efficiency Utilization of Solar Energy and Operation Control of Energy Storage System, Hubei University of Technology, Wuhan 430068, China.
None:
Accurate diagnosis of myocardial infarction (MI) holds significant clinical importance for public health systems. Deep learning-based ECG, classification and localization methods can automatically extract features, thereby overcoming the dependence on manual feature extraction in traditional methods. However, these methods still face challenges such as insufficient utilization of dynamic information in cardiac cycles, inadequate ability to capture both global and local features, and data imbalance. To address these issues, this paper proposes a ResNet-Transformer cascaded network (RTCN) to process time frequency features of ECG signals generated by the S-transform. First, the S-transform is applied to adaptively extract global time frequency features from the time frequency domain of ECG signals. Its scalable Gaussian window and high phase resolution can effectively capture the dynamic changes in cardiac cycles that traditional methods often fail to extract. Then, we develop an architecture that combines the Transformer attention mechanism with ResNet to extract multi-scale local features and global temporal dependencies collaboratively. This compensates for the existing deep learning models' insufficient ability to capture both global and local features simultaneously. To address the data imbalance problem, the Denoising Diffusion Probabilistic Model (DDPM) is applied to synthesize high-quality ECG samples for minority classes, increasing the inter-patient accuracy from 61.66% to 68.39%. Gradient-weighted Class Activation Mapping (Grad-CAM) visualization confirms that the model's attention areas are highly consistent with pathological features, verifying its clinical interpretability.