Related Experiment Video
Updated: Jul 5, 2026

Quantifying Learning in Young Infants: Tracking Leg Actions During a Discovery-learning Task
Published on: June 1, 2015
A Unified Learning and Evaluation Framework for Infant Cry-based Verification
Newborns communicate with the outside world primarily by crying. Infant cry-based verification can reduce the risk of mix-ups in hospital obstetrics. Recent studies have explored the potential of using infant cries for identity verification. Yet, model performance remains limited by training with variable-length clips and evaluating the complete audio recording from a single view. To this end, we propose a novel unified training and evaluation framework that uses fixed-length segments during training to ensure input consistency and incorporates a multi-view joint evaluation strategy by associating the audio recording with its local segments. Extensive experiments conducted on the public CryCeleb2023 dataset show that our framework leads to consistent improvements on different verification models. Specifically, the Equal Error Rate (EER) exhibited a reduction of 10.29% for the whisper-PMFA model, 6.63% for the X-Vector model, and 5.91% for the ECAPA-TDNN model. These results demonstrate the effectiveness of our fixed-length segment training and slice-based multi-view evaluation strategy in enhancing the model stability and evaluation accuracy, providing a more robust framework for newborn voice verification. The source code is released at https://github.com/contactless-healthcare/Unified-Infant-Cry-Verification.
Newborns communicate with the outside world primarily by crying. Infant cry-based verification can reduce the risk of mix-ups in hospital obstetrics. Recent studies have explored the potential of using infant cries for identity verification. Yet, model performance remains limited by training with variable-length clips and evaluating the complete audio recording from a single view. To this end, we propose a novel unified training and evaluation framework that uses fixed-length segments during training to ensure input consistency and incorporates a multi-view joint evaluation strategy by associating the audio recording with its local segments. Extensive experiments conducted on the public CryCeleb2023 dataset show that our framework leads to consistent improvements on different verification models. Specifically, the Equal Error Rate (EER) exhibited a reduction of 10.29% for the whisper-PMFA model, 6.63% for the X-Vector model, and 5.91% for the ECAPA-TDNN model. These results demonstrate the effectiveness of our fixed-length segment training and slice-based multi-view evaluation strategy in enhancing the model stability and evaluation accuracy, providing a more robust framework for newborn voice verification. The source code is released at https://github.com/contactless-healthcare/Unified-Infant-Cry-Verification.
More Related Videos
11:14A Novel Experimental and Analytical Approach to the Multimodal Neural Decoding of Intent During Social Interaction in Freely-behaving Human Infants
Published on: October 4, 2015
09:24Quantified Assessment of Infant's Gross Motor Abilities Using a Multisensor Wearable
Published on: May 17, 2024