Related Experiment Video
Updated: Apr 24, 2026

Recording Mouse Ultrasonic Vocalizations to Evaluate Social Communication
Published on: June 5, 2016
Spectrogram-derived graphs and inductive learning for multi-label avian vocalization detection in field recordings
1College of Engineering Trivandrum, APJ Abdul Kalam Technological University, Thiruvananthapuram, Kerala, India.
Abstract:
This paper presents a methodology that employs inductive spatial geometric deep learning networks to detect multiple avian vocalizations from field recordings. Initially, a graph is constructed from the Mel-spectrogram of each audio file using a trained deep convolutional neural network (Deep CNN). The extracted features are used to build a node-feature graph, which is then processed by two spatial inductive graph-based models: graph sample and aggregation (GraphSAGE) and the graph attention network (GAT), for multi-label classification. To enhance the robustness and generalization of the Deep CNN, SpecAugment is applied to generate additional Mel-spectrograms via data augmentation. The proposed framework is evaluated on the Xeno-canto bird sound database and compared against state-of-the-art methods. The results demonstrate that the proposed inductive spatial graph-based approach outperforms existing techniques, achieving macro F1-scores of 0.90 with GraphSAGE and 0.92 with GAT. We further replaced Deep CNN with AudioProtoPNet-20 and evaluated GAT on the Xeno-canto dataset, obtaining a macro F1-score of 0.93.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Sampling Methods: Overview
In analytical chemistry, the choice of...

