Related Experiment Video
Updated: Mar 24, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.4K
A misclassification-aware explainable hybrid CNN-vision transformer framework for radiographic weld inspection
Kumar Parmar1, Rituraj Jain1, P T Anitha2
1Department of Information Technology, Marwadi University, Rajkot, Gujarat, India.
Scientific Reports
|March 23, 2026
Summary
A new hybrid deep learning model combining Convolutional Neural Networks (CNN) and Vision Transformers (ViT) significantly improves weld defect identification accuracy. This intelligent framework enhances reliability and interpretability for Industry 5.0 inspection systems.
Area of Science:
- Materials Science and Engineering
- Artificial Intelligence
- Manufacturing Technology
Background:
- Accurate weld defect identification is crucial for structural integrity in safety-critical manufacturing.
- Existing deep learning methods struggle to differentiate visually similar weld defects, leading to costly misclassifications.
- There is a need for more reliable and interpretable weld inspection systems, particularly for Industry 5.0.
Purpose of the Study:
- To introduce a novel hybrid Convolutional Neural Network-Vision Transformer (CNN-ViT) architecture for intelligent weld inspection.
- To enhance the reliability and interpretability of automated weld defect detection systems.
- To improve the discrimination of visually similar defects and reduce misclassification rates.
Main Methods:
- A hybrid CNN-ViT architecture was developed and trained on the RIAWELC radiographic weld dataset.
- The proposed model was compared against a lightweight CNN baseline under identical training conditions.
- Explainability methods, including Grad-CAM and transformer self-attention maps, were used for misclassification analysis.
Main Results:
- The hybrid CNN-ViT achieved a superior accuracy of 98.56% compared to the CNN baseline (97.90%).
- A 31% relative reduction in the misclassification rate was observed with the hybrid model.
- Global contextual modeling in the CNN-ViT effectively distinguished visually similar defects like cracks and porosity.
Conclusions:
- The hybrid CNN-ViT framework offers improved accuracy and reduced misclassification for weld defect detection.
- The model provides transparent decision explanations, enhancing interpretability for industrial applications.
- This robust and explainable solution is well-suited for Industry 5.0-aligned intelligent weld inspection systems.