Related Experiment Video
Updated: Jun 28, 2026

12:08
From Voxels to Knowledge: A Practical Guide to the Segmentation of Complex Electron Microscopy 3D-Data
Published on: August 13, 2014
Cross-domain zero-shot semantic segmentation for unstructured environments via EVA-CLIP model, ensemble prompt
Nana Zhou1, Xianhua Zhao2, Fengjun Zhou1
1School of Computer Science, Shandong Xiehe University, Shandong, China.
Plos One
|June 26, 2026
Summary
This study introduces a novel framework for semantic segmentation in unstructured environments, enhancing zero-shot transferability for unmanned ground vehicles. The method significantly improves performance on benchmarks, rivaling supervised methods.
Area of Science:
- Computer Vision
- Robotics
- Artificial Intelligence
Background:
- Semantic segmentation is crucial for unmanned ground vehicles (UGVs) in unstructured environments for obstacle identification and path planning.
- Existing methods lack zero-shot transferability, requiring fine-tuning for novel scenarios and limiting adaptability.
Purpose of the Study:
- To develop a novel framework for robust zero-shot transfer in semantic segmentation for UGVs operating in unstructured domains.
- To enhance the adaptability and performance of semantic segmentation models without compromising zero-shot capabilities.
Main Methods:
- Leveraged the EVA-CLIP architecture for superior visual-linguistic alignment.
- Employed deep prompt tuning to adapt the EVA-CLIP image encoder for unstructured terrain features.
- Developed an ensemble prompt engineering scheme tailored for unstructured settings.
- Integrated global and local representations to optimize cross-modal alignment between text and images.
Main Results:
- Achieved significant improvements in mean Intersection over Union (mIoU) on the Robot Unstructured Ground Driving (RUGD) benchmark, ranging from 1.2% to 43.9%.
- Demonstrated cross-domain zero-shot performance on the Rellis-3D dataset comparable to supervised fine-tuning approaches.
- Showcased robust generalization capabilities to previously unseen semantic classes.
Conclusions:
- The proposed framework offers superior zero-shot transferability for semantic segmentation in unstructured environments.
- The methodology effectively enhances segmentation precision and adaptability for UGVs in complex terrains.
- This approach provides a promising alternative to traditional supervised fine-tuning for domain-specific segmentation tasks.