Related Experiment Video
Updated: May 25, 2025

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
348
Sequence-Aware Vision Transformer with Feature Fusion for Fault Diagnosis in Complex Industrial Processes.
Zhong Zhang1, Ming Xu2, Song Wang1
1School of Electrical and Information Engineering, Zhengzhou University, Zhengzhou 450001, China.
Entropy (Basel, Switzerland)
|February 26, 2025
Summary
This study introduces a novel Vision Transformer (ViT) model for industrial fault diagnosis, enhancing complex pattern recognition in high-dimensional time-series data. The global and local feature fusion approach improves accuracy and outperforms existing methods.
Area of Science:
- Engineering
- Computer Science
- Data Science
Background:
- Industrial fault diagnosis presents challenges with high-dimensional, long time-series data and complex couplings.
- Traditional methods and existing deep learning models like Vision Transformer (ViT) struggle with capturing both global temporal patterns and local information crucial for accurate diagnosis.
Purpose of the Study:
- To propose a novel global and local feature fusion sequence-aware ViT (GLF-ViT) for enhanced industrial fault diagnosis.
- To modify feature embedding in ViT to retain sampling point correlations and preserve local information.
Main Methods:
- Developed a GLF-ViT model that fuses global features from the classification token with local features from the encoder.
- Modified feature embedding to preserve sampling point correlations and local information.
- Conducted experiments analyzing data segment length, network depth, feature fusion, and attention head receptive field.
Main Results:
- Demonstrated that shallower encoder networks are more suitable for high-dimensional time-series fault diagnosis in complex industrial processes.
- The proposed GLF-ViT method significantly enhances complex fault diagnosis by effectively fusing global and local features.
- Outperformed state-of-the-art algorithms on the Tennessee Eastman (TE) dataset and a power transmission fault dataset.
Conclusions:
- The GLF-ViT model offers a significant advancement in industrial fault diagnosis, particularly for complex, high-dimensional time-series data.
- The findings suggest that a shallower network architecture combined with effective feature fusion is key to improving diagnostic accuracy.
- The approach shows strong generalizability and effectiveness across different industrial datasets.

