Video Experimental Relacionado
Updated: Sep 9, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.1K
EdgeVidCap: Un modelo ligero de subtítulos de video de canal espacial y doble rama para cámaras de borde de IoT
Lan Guo1, Xuyang Li1, Jinqiang Wang1
1School of Information Science and Engineering, Lanzhou University, Lanzhou 730000, China.
Sensors (Basel, Switzerland)
|August 28, 2025
Resumen
Este estudio presenta EdgeVidCap, un modelo ligero de subtítulos de video para cámaras de borde de IoT. Logra una comprensión eficiente del video y genera una descripción precisa en dispositivos con recursos limitados.
Área de la Ciencia:
- Inteligencia artificial
- Visión por computadora
- Internet de las cosas
Sus antecedentes:
- Las cámaras de borde inteligentes con computación de borde e integración de IoT permiten la comprensión de video local.
- Los modelos de subtítulos de video existentes son computacionalmente intensivos, lo que dificulta el despliegue en dispositivos IoT con recursos limitados.
Objetivo del estudio:
- Desarrollar un modelo ligero de subtítulos de video, EdgeVidCap, para un procesamiento eficiente en tiempo real en cámaras de borde de IoT.
- Para abordar las limitaciones de la alta complejidad computacional y los grandes recuentos de parámetros en las soluciones actuales de subtítulos de video.
Principales métodos:
- Propuso un módulo de Mamba de estado de atención sinérgica (SASM) que combina la atención del canal y los modelos de espacio de estado (SSM) para el modelado eficiente de las características espaciotemporales.
- Desarrolló un decodificador LSTM adaptativo guiado por la atención para la generación de subtítulos auto-regresivos con ponderación dinámica de características.
- Implementó un mecanismo de filtrado de marco simplificado para mejorar la eficiencia del procesamiento.
Principales resultados:
- EdgeVidCap demostró una mayor precisión en comparación con los métodos de subtítulos de video existentes en los conjuntos de datos MSR-VTT y MSVD.
- El modelo logró una mayor eficiencia de procesamiento debido a su filtrado de marco simplificado.
- Generó descripciones textuales más confiables después de la selección del marco.
Conclusiones:
- EdgeVidCap ofrece una solución efectiva para el subtítulo de video ligero en dispositivos de borde de IoT.
- El módulo SASM propuesto y el descodificador adaptativo LSTM contribuyen a una comprensión de vídeo eficiente y precisa.
- El modelo cumple con los requisitos de procesamiento en tiempo real para entornos de computación de borde con recursos limitados.
Palabras clave:
El IOTMecanismo de atenciónComputación de bordeRedes neuronales ligerasModelos espaciales de estadosubtítulos de vídeoMás Videos Relacionados
Videos de Conceptos Relacionados
Vision
55.3K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.3K
Light Acquisition
8.6K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
8.6K
Uniform Depth Channel Flow
144
Uniform depth channel flow keeps fluid depth consistent along channels such as irrigation canals. In natural channels, such as rivers, approximate uniform flow is often assumed. This condition occurs when the channel’s bottom slope matches the energy slope, balancing potential energy lost from gravity with head loss due to shear stress. This balance prevents depth changes along the channel length, resulting in a steady, uniform flow.Uniform flow in open channels with a constant cross-section...
144
Force Classification
1.6K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.6K
Multi-input and Multi-variable systems
149
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
149
Deconvolution
247
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
247

