Related Experiment Video
Updated: Jan 13, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.3K
Attention-Fusion-Based Two-Stream Vision Transformer for Heart Sound Classification
Kalpeshkumar Ranipa1, Wei-Ping Zhu1, M N S Swamy1
1Department of Electrical and Computer Engineering, Concordia University, Montreal, QC H3G 1M8, Canada.
Bioengineering (Basel, Switzerland)
|October 29, 2025
Summary
This study introduces an attention fusion-based two-stream Vision Transformer (AFTViT) for heart sound classification (HSC). The novel AFTViT architecture improves accuracy in diagnosing cardiovascular diseases.
Area of Science:
- Artificial Intelligence
- Biomedical Engineering
- Cardiology
Background:
- Heart sound classification (HSC) is crucial for cardiovascular disease diagnosis.
- Existing methods often use single-stream architectures, missing multi-resolution feature benefits.
- Current multi-stream approaches struggle with cross-modal interactions and information loss during fusion.
Purpose of the Study:
- To develop a novel attention fusion-based two-stream Vision Transformer (AFTViT) for enhanced heart sound classification.
- To effectively capture and integrate multi-resolution and cross-modal features in heart sound signals.
- To overcome limitations of conventional fusion methods in existing HSC architectures.
Main Methods:
- Proposed an AFTViT architecture utilizing two-dimensional mel-cepstral domain features.
- Employed a Vision Transformer (ViT)-based encoder for capturing long-range dependencies and multi-scale contextual information.
- Introduced a novel attention block for feature-level integration of cross-context features.
Main Results:
- The AFTViT architecture demonstrated superior performance compared to state-of-the-art CNN-based methods on PhysioNet2016 and PhysioNet2022 datasets.
- Achieved higher accuracy in heart sound classification tasks.
- The attention fusion mechanism effectively enhanced feature representation by integrating cross-contextual information.
Conclusions:
- The AFTViT framework shows significant potential for improving the accuracy of heart sound classification.
- This approach offers a promising tool for the early diagnosis of cardiovascular diseases.
- The study highlights the efficacy of attention-based fusion in multi-stream Vision Transformer architectures for biomedical signal processing.
Related Concept Videos
Heart Sounds
3.2K
Heart sounds are generated by the turbulence in blood flow due to the closing of heart valves. These sounds are best perceived slightly away from the valves, where the blood flow disseminates the sound.
Auscultation is the process of listening to these internal body sounds using a stethoscope. The heart produces four types of sounds, but only two—S1 and S2—can usually be heard with a stethoscope.
S1, also known as the "lub" sound, is caused by the closure of atrioventricular (A-V)...
Auscultation is the process of listening to these internal body sounds using a stethoscope. The heart produces four types of sounds, but only two—S1 and S2—can usually be heard with a stethoscope.
S1, also known as the "lub" sound, is caused by the closure of atrioventricular (A-V)...
3.2K
Parallel Processing
626
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
626
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K
