Related Experiment Video
Updated: Aug 11, 2026

06:37
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
5.3K
LargeAD: Large-Scale Cross-Sensor Data Pretraining for Autonomous Driving
IEEE Transactions on Pattern Analysis and Machine Intelligence
|October 2, 2025
Summary
LargeAD enables large-scale 3D pretraining for autonomous driving using vision foundation models (VFMs) and LiDAR data. This framework enhances 3D scene understanding by aligning 2D and 3D data for improved segmentation and detection.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Robotics
Background:
- Vision foundation models (VFMs) excel in 2D visual perception but are underexplored for 3D scene understanding in autonomous driving.
- Existing methods struggle with large-scale 3D data and cross-modal integration for real-world driving scenarios.
Purpose of the Study:
- To introduce LargeAD, a scalable framework for large-scale 3D pretraining using diverse driving datasets.
- To enhance 3D scene understanding by leveraging VFMs for cross-modal representation learning between 2D images and 3D LiDAR data.
Main Methods:
- Utilizing VFMs to generate semantically rich superpixels from 2D images.
- Aligning 2D superpixels with LiDAR point clouds for contrastive sample generation.
- Implementing VFM-assisted contrastive learning, superpoint temporal consistency, and multi-source data pretraining.
Main Results:
- Achieved significant performance gains over state-of-the-art methods in linear probing and fine-tuning.
- Demonstrated superior results in LiDAR-based segmentation and object detection tasks.
- Validated adaptability, efficiency, and robustness across 11 large-scale multi-sensor datasets.
Conclusions:
- LargeAD provides a versatile and scalable solution for 3D pretraining in autonomous driving.
- The framework effectively enhances semantic consistency and representation learning across 2D and 3D modalities.
- LargeAD shows strong potential for real-world autonomous driving applications requiring robust 3D scene understanding.