Related Experiment Video
Updated: Jun 14, 2025

Experience is Instrumental in Tuning a Link Between Language and Cognition: Evidence from 6- to 7- Month-Old Infants' Object Categorization
Published on: April 19, 2017
Multi-Task Learning for Audio-Based Infant Cry Detection and Reasoning
Insights
This study introduces a new Infant Cry Detection and Reasoning (ICDR) model to improve infant cry analysis. The ICDR model enhances generalization by using multi-task learning and a novel contrastive mixture of experts approach.
Area of Science:
- Machine Learning
- Infant Health Monitoring
- Signal Processing
Background:
- Infant cry analysis is vital for understanding infant well-being, but limited datasets and individual voice variations hinder model performance.
- Existing models struggle with generalization due to data scarcity and subject-specific acoustic differences in infant cries.
Purpose of the Study:
- To develop a robust multi-task model for Infant Cry Detection and Reasoning (ICDR) that overcomes data limitations and subject variability.
- To improve the accuracy and generalization capabilities of AI models in interpreting infant vocalizations.
Main Methods:
- Proposed a multi-task learning framework leveraging two datasets to increase data diversity.
- Introduced an efficient attention module for inter-task feature enrichment.
- Implemented an intra-task contrastive mixture of experts (CMoE) module to reduce subject variance and enhance representation consistency.
Main Results:
- The ICDR model demonstrated superior performance compared to state-of-the-art methods in infant cry detection and reasoning.
- Achieved significant improvements in F1-score (2-9%) across extensive cross-subject experiments.
- Validated the effectiveness of multi-task learning and the CMoE module in enhancing model generalization.
Conclusions:
- Multi-task learning with inter-task attention and intra-task CMoE significantly boosts the generalization ability of infant cry analysis models.
- The proposed ICDR model offers a promising solution for more reliable infant cry interpretation in real-world applications.
Abstract:
Infant cry is a crucial indicator that offers valuable insights into their physical and mental conditions, such as hunger and pain. However, the scarcity of infant cry datasets hinders the model's generalization in real-life scenarios. The varying voiceprint characteristics among infants further exacerbate this challenge, deteriorating the model's performance on unseen infants. To this end, we propose a multi-task model for Infant Cry Detection and Reasoning (ICDR). It leverages datasets from two tasks to enrich data diversity and introduces an efficient attention module to achieve inter-task feature supplementarity. To mitigate the impact of subject differences, ICDR introduces an intra-task contrastive mixture of experts (CMoE) module that adaptively allocates experts to reduce subject variance and applies contrastive learning to enhance the representation consistency of samples from different infants in the same state. Extensive cross-subject experiments show that ICDR outperforms the state-of-the-art models in infant cry detection and reasoning, with an improvement of 2-9% in the F1-score. This demonstrates that multi-task learning effectively enhances the model's generalization ability by inter-task attention and intra-task CMoE.

