A Scene Adaption Framework for Infant Cry Detection in Obstetrics
Insights
This study introduces a Scene Adaptation Framework (SAF) to improve infant cry detection in clinical settings. SAF enhances model performance by adapting to new environments using acoustic principles and unsupervised learning, boosting F1-scores by 30%.
Area of Science:
- Medical acoustics
- Machine learning in healthcare
- Infant health monitoring
Background:
- Infant cry analysis offers critical clinical insights for medical decisions, particularly in obstetrics.
- Current infant cry detection models struggle in real clinical settings due to limited, specific training data.
- Developing robust cry detection is essential for timely and accurate caregiver interventions.
Purpose of the Study:
- To propose a Scene Adaptation Framework (SAF) for rapidly adapting infant cry detection models to new clinical environments.
- To address the challenge of limited training data in real-world clinical scenarios for infant cry detection.
- To enhance the F1-score performance of infant cry detection classifiers in obstetrics.
Main Methods:
- SAF employs a two-stage learning process: imitating clinical sounds using public datasets via acoustic principles and unsupervised adaptation using mutual learning.
- The first stage leverages the additive nature of audio signal mixtures to simulate clinical acoustics.
- The second stage uses mutual learning to extract shared infant cry features between clinical and public datasets.
Main Results:
- A clinical trial in Obstetrics with 200 infants demonstrated significant improvements in cry detection.
- Four tested classifiers showed nearly a 30% increase in F1-score when utilizing the SAF.
- SAF achieved performance comparable to supervised learning models trained directly on target clinical data.
Conclusions:
- The Scene Adaptation Framework (SAF) is an effective, plug-and-play solution for enhancing infant cry detection in novel clinical settings.
- SAF significantly improves classifier performance, overcoming limitations posed by scarce clinical training data.
- The framework demonstrates the potential for broader application in adapting AI models to specific healthcare environments.
Abstract:
Infant cry provides useful clinical insights for caregivers to make appropriate medical decisions, such as in obstetrics. However, robust infant cry detection in real clinical settings (e.g. obstetrics) is still challenging due to the limited training data in this scenario. In this paper, we propose a scene adaption framework (SAF) including two different learning stages that can quickly adapt the cry detection model to a new environment. The first stage uses the acoustic principle that mixture sources in audio signals are approximately additive to imitate the sounds in clinical settings using public datasets. The second stage utilizes mutual learning to mine the shared characteristics of infant cry between the clinical setting and public dataset to adapt the scene in an unsupervised manner. The clinical trial was conducted in Obstetrics, where the crying audios from 200 infants were collected. The experimented four classifiers used for infant cry detection have nearly 30% improvement on the F1-score by using SAF, which achieves similar performance as the supervised learning based on the target setting. SAF is demonstrated to be an effective plug- and-play tool for improving infant cry detection in new clinical settings. Our code is available at https://github.com/contactless-healthcare/Scene-Adaption-for-Infant-Cry-Detection.


