Related Experiment Video
Updated: Feb 17, 2026

Author Spotlight: UAV Remote Sensing for Efficient Invasive Plant Biomass Estimation
Published on: February 9, 2024
Vision-language models for zero-shot weed detection and visual reasoning in UAV-based precision agriculture
Muhammad Fahad Nasir1, Mobeen Ur Rehman2, Irfan Hussain1
1Khalifa University Center for Autonomous Robotic Systems, Khalifa University, Abu Dhabi, United Arab Emirates.
Vision-language models (VLMs) show promise for precision weed management in agriculture. Gemini Flash 2.5 demonstrated strong zero-shot performance and interpretability, offering a low-annotation alternative to traditional deep learning methods for drone imagery analysis.
Area of Science:
- Agricultural Science
- Computer Vision
- Artificial Intelligence
Background:
- Weeds significantly reduce row-crop yields, necessitating effective management strategies.
- Current deep learning methods for weed detection using UAV imagery require extensive data annotation and exhibit poor generalization.
- Limited interpretability of existing models hinders trust and adoption in agricultural applications.
Purpose of the Study:
- To evaluate the efficacy of modern vision-language models (VLMs) in a zero-shot setting for weed detection and management in soybean fields.
- To assess the capabilities of VLMs in identifying weed presence, spatial localization, crop type, and growth stage.
- To introduce and validate Error-Probing Prompting (EPP) for enhancing VLM interpretability and self-correction.
Main Methods:
- Six commercial VLMs (ChatGPT-4.1, ChatGPT-4o, Gemini Flash 2.5, Gemini Flash Lite 2.5, LLaMA-4 Scout, LLaMA-4 Maverick) were evaluated using drone imagery from soybean fields.
- A unified prompt was designed to elicit weed presence, spatial localization, reasoning, crop growth stage, and crop type.
- Error-Probing Prompting (EPP) was introduced as a counterfactual analysis method to test model robustness and self-correction capabilities.
- Interpretability was quantified using expert-rated scores for Grounding, Specificity, Plausibility, Non-Hallucination, and Actionability.
Main Results:
- Gemini Flash 2.5 exhibited the most consistent zero-shot performance and highest interpretability.
- ChatGPT-4.1 showed strong reasoning but lower detection accuracy, while ChatGPT-4o provided a balanced performance.
- LLaMA-4 variants demonstrated limitations in localization and specificity.
- Gemini Flash Lite 2.5 was efficient but lacked robustness under EPP stress tests, indicating brittle reasoning.
- Interpretability scores correlated positively with spatial correctness, as indicated by visual grounding and text-to-region overlap metrics.
Conclusions:
- Vision-language models offer a promising low-annotation approach for precision weed management.
- Model reliability in field deployment is better predicted by explainability and adaptability than by scale alone.
- Gemini Flash 2.5 emerged as a top performer, highlighting the potential of VLMs for practical agricultural applications.
- Further research into VLM explainability and feedback-driven adaptability is crucial for advancing automated agricultural systems.
Related Concept Videos
Depth Perception and Spatial Vision
Light Acquisition
Vision
Application of Linearization and Approximation
