Related Experiment Video
Updated: May 26, 2026

Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
Published on: July 2, 2013
AQFormer: severity-aware transformer with aphasia-specific CAM for spoken keyword classification in aphasic speech
Gowri Prasood Usha1, John Sahaya Rani Alex1
1School of Electronics Engineering, Vellore Institute of Technology, Chennai, Tamil Nadu, India.
Abstract:
Language-driven speech output in individuals with aphasia shows considerable variability, including phonological errors and pauses during word searches. This makes it difficult to use traditional keyword classification systems and further reduces trust in deep neural models, complicating their application in clinical settings. This paper introduces AQFormer, a severity-aware transformer architecture designed to classify spoken keywords in aphasic speech, and A-CAM, a dual-stream attribute framework aimed at assisting individuals with aphasic impairments. AQFormer generates acoustic representations that are severity-adaptive by integrating patient-level Aphasia Quotient (AQ) scores through Feature-wise Linear Modulation (FiLM) and A-CAM. A-CAM consists of two main components: (i) a branch that influences WavLM convolutional features, a prediction-focused one, and (ii) a multimodal aphasia filter that captures pauses, phoneme variations, and interruptions at word boundaries, an impairment-focused branch. We introduce an adaptive perturbation and dual-filtering gradient scheme that enforces non-negative, mask-consistent attributions over time-frequency regions. Experiments utilizing a subset of AphasiaBank keywords (93 speakers, 960 recordings; training set expanded to 5,138) with rigorous speaker-disjoint evaluation indicate that AQFormer achieves approximately 96.61% accuracy (F1 = 96.8%) on previously unseen speakers. A-CAM consistently outperforms several Grad-CAM variants when deletion/insertion AUPC and ADCC metrics are employed. This results in stable, sparse explanations that reflect how aphasia is usually caused: Discriminates correct from incorrect productions with Cohen's d = 2.05 (a massive effect size) and spatial localization of error regions with Intersection over Union (IoU) of 0.461 against phoneme boundaries. Montreal Forced Aligner meets the quantitative validation criteria for the aphasia filter. The impairment-focused A-CAM maps achieve an IoU of 0.712 against detected error regions, with a severity correlation that doubles from rho = -0.374 (base) to rho = -0.754 (filter-gated). By tightly coupling severity-aware modelling with aphasia-informed attributes, the proposed framework advances explainable learning systems for aphasia-affected speech without needing clinician-labelled training targets.
Related Concept Videos
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
