Related Experiment Video
Updated: Feb 14, 2026

A View of Their Own: Capturing the Egocentric View of Infants and Toddlers with Head-Mounted Cameras
Published on: October 5, 2018
Screen Detection from Egocentric Image Streams Leveraging Multi-View Vision Language Model
Xueshen Li1, Sen Shen2, Xinlong Hou1
1Department of Biomedical Engineering, Stevens Institute of Technology.
Abstract:
Accurately monitoring the screen exposure of young children is important for research related to screen use, such as childhood obesity, physical activity, and social interaction. Most existing studies rely upon self-report or manual measures from bulky wearable sensors, thus lacking efficiency and accuracy in capturing quantitative screen exposure data. In this work, we developed a novel screen detection framework that utilizes egocentric images from a wearable sensor, named the screen time tracker (STT), and a vision language model (VLM). In particular, we devised a multi-view VLM that takes multiple views from egocentric image streams and interprets screen exposure dynamically. We validated our approach by using a dataset of children's free-living activities, demonstrating significant improvement over existing methods in conventional vision language models and object detection models. The combination of a vision language model and a lightweight hardware design provides a novel solution in screen detection for children. The proposed framework has great potential to benefit children's behavioral study. The code is available at https://github.com/YGanLab/MV-VLM.
Related Concept Videos
Vision
Stream Function
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Color Vision
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

