Related Experiment Video
Updated: May 8, 2026

Automated Interactive Video Playback for Studies of Animal Communication
Published on: February 9, 2011
Artificial intelligence-powered 3D analysis of video-based caregiver-child interactions
Zhenzhen Weng1, Laura Bravo-Sánchez2, Zeyu Wang2
1Institute for Computational & Mathematical Engineering, Stanford University, Stanford, CA, USA.
Abstract:
We introduce HARMONI, a three-dimensional (3D) computer vision and audio processing method for analyzing caregiver-child behavior and interaction from observational videos. HARMONI operates at subsecond resolution, estimating 3D mesh representations and spatial interactions of humans, and adapts to challenging natural environments using an environment-targeted synthetic data generation module. Deployed on 500 hours from the SEEDLingS dataset, HARMONI generates detailed quantitative measurements of 3D human behavior previously unattainable through manual efforts or 2D methods. HARMONI identifies longitudinal trends in child-caregiver interaction, including child movement, body pose, dyadic touch, visibility, and conversational turns. The integrated visual and audio analysis further reveals multimodal trends, including associations between child conversational turns and movement. Open-sourced for large-scale analysis, HARMONI facilitates advancements in human development research. HARMONI achieves 63 to 80% consistency on key attributes with human annotators on SEEDLingS and 84 to 93% consistency on videos taken from a laboratory setting while achieving >100 times savings in time.
More Related Videos
10:11Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism
Published on: December 14, 2012
08:08Author Spotlight: Capturing Infant-Caregiver Interactions Through Synchronized Multimodal Data Collection
Published on: May 31, 2024