Related Experiment Video
Updated: Feb 10, 2026

Development of an Audio-based Virtual Gaming Environment to Assist with Navigation Skills in the Blind
Published on: March 27, 2013
Towards Robust Family-Infant Audio Analysis Based on Unsupervised Pretraining of Wav2vec 2.0 on Large-Scale Unlabeled
Jialu Li1,2, Mark Hasegawa-Johnson1,2, Nancy L McElwain2,3
1Department of Electrical and Computer Engineering, University of Illinois.
This study introduces LittleBeats (LB), an infant wearable device, using wav2vec 2.0 (W2V2) for family audio analysis. W2V2 pretraining on home recordings improved speaker diarization and vocalization classification for infants.
Area of Science:
- Speech processing
- Machine learning
- Infant behavior analysis
Background:
- Automatic family audio analysis requires robust speech processing models.
- Previous methods utilized general-purpose embeddings and supervised learning.
- Infant wearable devices offer new avenues for data collection.
Purpose of the Study:
- To enhance the audio analysis capabilities of the LittleBeats (LB) infant wearable device.
- To investigate the effectiveness of wav2vec 2.0 (W2V2) pretraining on home recordings for family audio representation.
- To compare W2V2 performance with limited labeled data against large unlabeled datasets.
Main Methods:
- Pretraining wav2vec 2.0 (W2V2) on 1k-hour of unlabeled home recordings from the LB device.
- Fine-tuning W2V2 for speaker diarization (SD) and vocalization classification (VC) tasks.
- Utilizing SpecAugment and environmental speech corruptions to improve model robustness.
Main Results:
- W2V2 pretrained on 1k-hour home recordings outperformed models pretrained on 52k-hour general audio for LB data.
- Achieved a 12% relative gain in speaker diarization (SD) and a moderate boost in vocalization classification (VC).
- External unlabeled and labeled data further enhanced W2V2 pretraining and fine-tuning.
Conclusions:
- W2V2 pretraining on domain-specific home recordings is effective for infant audio analysis.
- Limited, relevant data can yield superior results compared to large, general datasets.
- The LittleBeats device and W2V2 model offer a promising approach for family audio analysis.
More Related Videos
13:08Measurement of Fronto-limbic Activity Using an Emotional Oddball Task in Children with Familial High Risk for Schizophrenia
Published on: December 2, 2015
07:31Investigating the Effect of Visual Imagery and Learning Shape-Audio Regularities on Bouba and Kiki
Published on: September 13, 2019
Related Concept Videos
Protein Families
Protein Families
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Gene Families
Family Therapy
Strategic Family Therapy
Strategic family therapy emphasizes resolving communication barriers and improving problem-solving abilities...
Sources of Self-Esteem I: Family Experience