Related Experiment Video
Updated: Feb 14, 2026

03:56
A View of Their Own: Capturing the Egocentric View of Infants and Toddlers with Head-Mounted Cameras
Published on: October 5, 2018
8.0K
Screen Detection from Egocentric Image Streams Leveraging Multi-View Vision Language Model
Xueshen Li1, Sen Shen2, Xinlong Hou1
1Department of Biomedical Engineering, Stevens Institute of Technology.
Summary
Researchers developed a new screen detection system for children using a wearable camera and AI. This method accurately tracks screen time, improving child behavior studies and screen use research.
Area of Science:
- Child development
- Human-computer interaction
- Wearable technology
Background:
- Accurate monitoring of children's screen time is crucial for research on health and social behaviors.
- Current methods like self-reports and bulky sensors lack accuracy and efficiency.
- Existing technologies struggle to capture quantitative screen exposure data effectively.
Purpose of the Study:
- To develop an efficient and accurate framework for monitoring children's screen exposure.
- To introduce a novel approach combining wearable sensors and advanced AI for screen time tracking.
- To improve the methodologies used in studying screen use and its impact on child development.
Main Methods:
- Developed a novel screen detection framework using egocentric images from a wearable sensor called the screen time tracker (STT).
- Utilized a multi-view vision language model (VLM) to dynamically interpret screen exposure from multiple image streams.
- Validated the framework using a dataset of children's free-living activities.
Main Results:
- The proposed multi-view VLM framework demonstrated significant improvements over conventional vision language models and object detection models.
- The system achieved higher accuracy and efficiency in capturing quantitative screen exposure data.
- The lightweight hardware design of the STT combined with the VLM offers a practical solution.
Conclusions:
- The novel screen detection framework provides an accurate and efficient method for monitoring children's screen time.
- This technology has significant potential to advance research in child behavior, screen use, and related health outcomes.
- The combination of wearable sensors and VLM offers a promising direction for future child behavior studies.
Related Concept Videos
Vision
60.3K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
60.3K
Stream Function
2.1K
In two-dimensional incompressible fluid flow, the continuity equation is essential for ensuring mass conservation, meaning that any change in fluid entering or exiting a region is balanced by a corresponding change elsewhere. For incompressible flow, where density remains constant, this requirement simplifies to the condition that the divergence of the velocity field must be zero. Mathematically, this is expressed as,
2.1K
Language
921
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
921
Color Vision
1.5K
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
1.5K
Components of Language
831
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
831
Language Development
939
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
939

