Related Experiment Videos
CMS-Attack: A Structured Cross-Modal Search Attack for Robustness Evaluation of LiDAR-Camera Fusion Detectors
Minzhou Wang1, Yaoguang Cao2, Shichun Yang2
1School of Vehicle and Energy, Yanshan University, No. 438 West Hebei Avenue, Haigang District, Qinhuangdao 066004, China.
Abstract:
LiDAR-camera fusion is widely used for 3D perception in intelligent connected vehicles, but a clean camera branch does not necessarily compensate for structured LiDAR corruption. We propose CMS-Attack, a cross-modal search framework in which only the LiDAR point cloud is perturbed while the camera input remains unchanged; here, "cross-modal" denotes that a single-modality LiDAR perturbation propagates through the LiDAR-camera fusion process and disrupts the multimodal detector, rather than simultaneous perturbation of both modalities. The framework has the following two access-dependent routes: the gray-box route contains FB-CMS, which uses camera-BEV, LiDAR-BEV, fused-BEV, and detection-head responses to construct a target-aware prior and prune an over-complete candidate pool, and Adaptive CMS, which substitutes architecture-specific intermediate responses; the decision-only black-box route contains FC-CMS, which refines candidates solely from display-level target states. On the nuScenes validation split, FB-CMS reduced matched target confidence from 0.80 to 0.03 under 140 injected points, corresponding to a 96.1% relative drop and 100% ASR@0.3. The ten-query FC-CMS achieved 58.97% ASR@0.3, compared with 6.17% for random frustum spoofing. On the query-based FUTR3D detector, Adaptive CMS reduced the mean matched-target score from 0.642 to 0.058, corresponding to a 90.97% relative reduction and 93.42% ASR@0.3. Intermediate camera-, LiDAR-, and fused-BEV region energies changed by less than 0.8% despite target suppression, indicating disruption at the fusion-decision stage (defined here as the decoder/detection-head and post-processing path from fused representations to final object predictions) rather than a collapse of BEV feature magnitude. These results show that structured LiDAR fabrication is substantially more disruptive than information removal and that generic outlier filtering incurs a robustness-accuracy tradeoff.
Related Concept Videos
Confocal Fluorescence Microscopy
Light Acquisition