Related Experiment Video
Updated: Jul 21, 2026

11:41
Evaluation of an Exclusive Spur Dike U-Turn Design with Radar-Collected Data and Simulation
Published on: February 1, 2020
20.3K
A Vision-Language Model-Based Traffic Sign Detection Method for High-Resolution Drone Images: A Case Study in Guyuan,
Jianqun Yao1, Jinming Li1, Yuxuan Li1
1CCCC Infrastructure Maintenance Group Co., Ltd., Beijing 100011, China.
Sensors (Basel, Switzerland)
|September 14, 2024
Summary
This study introduces Vision-Language Models (VLMs) for efficient traffic sign detection using drone imagery. This approach eliminates the need for manual annotations, significantly reducing training costs and time.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Transportation Engineering
Background:
- Traffic signs are crucial for transportation systems, with drones increasingly used for monitoring.
- Current traffic sign detection methods rely heavily on time-consuming image annotations.
- Developing large, diverse datasets for traffic sign monitoring is a significant challenge.
Purpose of the Study:
- To explore the application of Vision-Language Models (VLMs) for traffic sign detection using drone images.
- To develop a method that bypasses the need for discrete image labels, enabling rapid deployment.
- To reduce the cost and time associated with creating annotated datasets for traffic sign monitoring.
Main Methods:
- Compiled a keyword dictionary for traffic signs, incorporating Chinese national standards for shape and color.
- Utilized Bootstrapping Language-image Pretraining v2 (BLIPv2) to generate text descriptions from representative images.
- Employed a Contrastive Language-image Pretraining (CLIP) framework for characterizing drone images and text descriptions.
- Predicted traffic sign categories based on similarity between visual features and word embeddings using cosine distance and softmax.
Main Results:
- The proposed VLM-based method demonstrated acceptable prediction accuracy in practical applications using drone images from Guyuan, China.
- Experiments on two public datasets confirmed the method's effectiveness.
- The approach achieved a low training cost compared to traditional annotation-reliant methods.
Conclusions:
- Vision-Language Models offer a viable and efficient alternative for traffic sign detection from drone imagery.
- The developed method significantly reduces the reliance on manual annotations, lowering dataset creation costs.
- This approach holds promise for scalable and rapid deployment in traffic sign monitoring systems.

