Related Experiment Video
Updated: Aug 14, 2026

End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA
Abhishek Gupta1, Ajmery Sultana1
1Faculty of Computer Science and Technology, Algoma University, Brampton, ON L6V 1A3, Canada.
This study introduces the Latent Model-Joint Embedding Predictive Architecture (LM-JEPA), a resource-efficient framework for connected and autonomous vehicles. LM-JEPA enhances collaborative perception by integrating latent predictive learning with multi-modal reasoning for improved scene understanding.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Robotics
Background:
- Autonomous driving demands efficient scene understanding under strict resource constraints (latency, energy, communication).
- Large Language Models (LLMs) and Vision-Language Models (VLMs) face limitations in edge deployments due to computational demands.
- Existing latent-space approaches often focus on single-modal perception, lacking integrated multi-modal reasoning for collaborative tasks.
Purpose of the Study:
- To present the Latent Model-Joint Embedding Predictive Architecture (LM-JEPA) for resource-efficient collaborative perception in connected and autonomous vehicles.
- To enable efficient perception and reasoning in edge environments by encoding heterogeneous sensor data into a unified latent space.
- To integrate multi-modal latent reasoning and adaptive sensor fusion for enhanced performance under resource constraints.
Main Methods:
- Encoding heterogeneous inputs (camera, LiDAR, radar, map) into a unified latent space using a joint embedding predictive architecture.
- Implementing a context-adaptive multi-modal fusion mechanism for dynamic weighting of sensor and model contributions.
- Integrating a lightweight VLM with an edge-assisted pipeline for real-time inference, adaptive offloading, and latent-space reasoning for cooperative decision-making.
Main Results:
- LM-JEPA improves perception accuracy by 5% and reduces latency by approximately 7% compared to LLM and VLM baselines.
- Achieved up to 25% improvement in scene understanding and 20% higher intersection success rates.
- Demonstrated improved highway merging capabilities and approximately 15% reduction in transmitted model parameters.
Conclusions:
- LM-JEPA offers a practical and resource-efficient solution for collaborative perception in autonomous vehicles operating under edge constraints.
- The framework effectively integrates multi-modal data and lightweight reasoning for enhanced scene understanding and decision-making.
- LM-JEPA represents a significant advancement in enabling robust and efficient autonomous driving systems.
Related Concept Videos
Light Acquisition
Automatic Processing and Automatic Social Behavior
Depth Perception and Spatial Vision
Perception
Bottom-up processing begins at the sensory level, where receptors detect external environmental stimuli. These could include the tactile sensation of...
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...