Related Experiment Videos
Multibranch Attention and Fusion Network for Voice-Based Autism Detection Using Heterogeneous Audio Embeddings
Mouad El Omari1, Younes El Belghiti1, Hanae Belmajdoub1,2
1LRIT Laboratory, Faculty of Sciences, Mohammed V University, Rabat, Morocco.
Abstract:
Voice-based analysis is attracting growing interest as a noninvasive means of identifying early markers of autism spectrum disorder (ASD). While pretrained audio models such as YAMNet and VGGish provide complementary views of children's speech, most existing studies rely on a single representation and do not explore how these embeddings may be combined in a structured manner. This work introduces the multibranch attention and fusion network (MBAFNet), an architecture designed to make fuller use of heterogeneous embeddings by processing each stream through its own convolutional encoder, extracting temporal cues at multiple scales, and modeling cross-representation interactions through a self-attention layer. A gating mechanism then regulates the relative contribution of each embedding before classification. Experiments conducted on the CASD-SC corpus under a subject-disjoint five-fold evaluation protocol show that MBAFNet achieves 94.17% accuracy, outperforming all evaluated baselines and previously reported state-of-the-art approaches. These findings indicate that carefully designed selective fusion can reveal complementary information contained in pretrained embeddings and supports the development of more robust speech-based ASD assessment frameworks, while broader validation remains necessary before screening-oriented use.