Infant Cry Signal Diagnostic System Using Deep Learning and Fused Features

Yara Zayed1, Ahmad Hasasneh1, Chakib Tadj2

  • 1Department of Natural, Engineering and Technology Sciences, Faculty of Graduate Studies, Arab American University, Ramallah P.O. Box 240, Palestine.

Insights

Diagnosing infant conditions like sepsis and respiratory distress syndrome (RDS) is challenging. This study developed a deep learning system analyzing infant cry audio signals (CAS) to accurately detect these critical conditions with 97.50% accuracy.

Area of Science:

  • Medical Diagnostics
  • Bioacoustics
  • Machine Learning

Background:

  • Infants cannot verbalize symptoms, making early diagnosis difficult.
  • Infant crying is a primary communication method for distress and needs.
  • Accurate diagnosis of conditions like neonatal respiratory distress syndrome (RDS) and sepsis is vital due to high mortality rates.

Purpose of the Study:

  • To develop and evaluate a medical diagnostic system for interpreting infant cry audio signals (CAS).
  • To improve the accuracy and classification rate of diagnosing infant pathologies using cry analysis.
  • To leverage deep learning algorithms and fused audio features for enhanced diagnostic capabilities.

Main Methods:

  • Utilized a dataset of labeled infant cry audio signals, including those with RDS, sepsis, and healthy cries.
  • Extracted audio features: harmonic ratio (HR), Gammatone frequency cepstral coefficients (GFCCs), and spectrograms via a pre-trained CNN.
  • Fused these features and applied machine learning models (RF, SVM, DNN), with a focus on deep learning for feature extraction and fusion.

Main Results:

  • The system achieved a highest accuracy of 97.50% using fused spectrogram, HR, and GFCC features processed through a deep learning model.
  • Feature fusion through the learning process, particularly with spectrograms, significantly improved classification compared to simple concatenation.
  • The deep learning approach effectively extracted sparsely represented features, enhancing the separation between different infant pathologies.

Conclusions:

  • Fusing diverse audio features, especially spectrograms, within a deep learning framework is crucial for accurate infant cry analysis.
  • The proposed system demonstrates a promising, non-invasive method for the early diagnosis of critical infant medical conditions.
  • This approach outperforms previous benchmarks by enabling multi-classification of pathologies using advanced feature engineering and deep learning.