UAVベースの精密農業におけるゼロショット雑草検出と視覚推論のための視覚言語モデル
Muhammad Fahad Nasir1, Mobeen Ur Rehman2, Irfan Hussain1
1Khalifa University Center for Autonomous Robotic Systems, Khalifa University, Abu Dhabi, United Arab Emirates.
Frontiers in plant science
|February 16, 2026
まとめ
視覚言語モデル (Vision-language models,VLMs) は,農業における precision 雑草管理の有望性を示しています. Gemini Flash 2.5は,強力なゼロショット性能と解釈性を実証し,ドローンの画像分析のための従来のディープラーニング方法の代替案として低アノテーションを提供しました.
科学分野:
- 農業科学 農業科学とは
- コンピュータビジョン コンピュータビジョン
- 人工知能 (AI) とは,人工知能 (AI) のことです.
背景:
- 雑草は,列作物の収穫量を大幅に低下させ,効果的な管理戦略を必要とします.
- UAV画像を使用した雑草検出のための現在のディープラーニング方法は,広範なデータアノテーションを必要とし,一般的な表現が悪い.
- 既存のモデルの解釈が限られていることは,農業の応用に対する信頼と採用を妨げています.
研究 の 目的:
- 豆畑における雑草の検出と管理のためのゼロショット設定における近代的な視覚言語モデル (VLMs) の有効性を評価する.
- 雑草の存在,空間的位置,作物の種類,成長段階を特定するVLMの能力を評価する.
- VLMの解釈性と自己訂正性を向上させるため,エラー検定プロンプト (EPP) を導入し,検証する.
主な方法:
- 6機の商用VLM (ChatGPT-4.1,ChatGPT-4o,Gemini Flash 2.5,Gemini Flash Lite 2.5,LLaMA-4 Scout,LLaMA-4 Maverick) は,大豆畑からのドローン画像を用いて評価された.
- 統一されたプロンプトは,雑草の存在,空間的局所,推論,作物の成長段階,作物の種類を誘発するように設計されました.
- エラーテストプロンプト (EPP) は,モデルの堅強さと自己訂正能力をテストするための対事実分析方法として導入されました.
- 解釈可能性は,根拠,特異性,妥当性,非幻覚性,可動性に関する専門家評価のスコアを使用して定量化されました.
主要な成果:
- Gemini Flash 2.5は,最も一貫したゼロショット性能と最も高い解釈性を示しました.
- ChatGPT-4.1は強力な推論を示したが,検出精度は低く,ChatGPT-4oはバランスの取れたパフォーマンスを提供した.
- LLaMA-4の変種は,局所化と特異性の限界を示した.
- Gemini Flash Lite 2.5は効率的でしたが,EPPのストレステストでは頑丈性がなく,脆い推論を示しています.
- 解釈可能性のスコアは,視覚的な接地とテキストから地域への重複メトリックによって示されるように,空間的正しさと正に相関していました.
結論:
- 視覚言語モデルは,精密な雑草管理のための有望な低アノテーションのアプローチを提供します.
- フィールド展開におけるモデルの信頼性は,スケールだけで予測するよりも,説明性と適応性によってよりよく予測されます.
- Gemini Flash 2.5は,実用的な農業アプリケーションのためのVLMの可能性を強調して,トップパフォーマンスとして浮上しました.
- VLMの説明性とフィードバックに基づく適応性に関するさらなる研究は,自動化された農業システムの進歩に不可欠です.
関連する概念動画
Depth Perception and Spatial Vision
2.1K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.1K
Light Acquisition
9.7K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
9.7K
Vision
60.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
60.4K
Application of Linearization and Approximation
105
A drone flying through complex terrain often relies on more than one sensing method to estimate small changes in altitude. Along with direct measurements, air pressure provides a useful indirect indicator of vertical movement. Atmospheric pressure decreases as altitude increases, and this relationship is commonly described using an exponential model. Although accurate, converting pressure measurements into altitude values requires calculations that are too complex to perform repeatedly during...
105

