Related Experiment Video
Updated: Aug 12, 2026

08:47
Computer Vision-Based Biomass Estimation for Invasive Plants
Published on: February 9, 2024
A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
Syed Nazmus Sakib1, Nafiul Haque1, Mohammad Zabed Hossain2
1Department of Robotics and Mechatronics Engineering, University of Dhaka, Dhaka, Bangladesh.
Scientific Data
|August 10, 2026
Summary
A new visual question answering (VQA) dataset, PlantExpertVQA, enables advanced AI for plant disease diagnosis. Fine-tuning a small AI model on this dataset significantly improves agricultural decision-making capabilities.
Area of Science:
- Agricultural Science
- Computer Vision
- Artificial Intelligence
Background:
- Existing plant disease datasets primarily support classification and detection tasks.
- This limits the application of vision-language models in interactive, reasoning-based plant disease diagnosis.
- There is a need for comprehensive datasets to train AI for complex agricultural decision-making.
Purpose of the Study:
- To introduce PlantExpertVQA, a large-scale visual question answering dataset for agricultural decision-making.
- To facilitate the development of advanced vision-language models for plant disease diagnosis.
- To benchmark current AI models and demonstrate effective domain adaptation strategies.
Main Methods:
- Compiled 765,186 question-answer pairs from 45 open-source datasets, including PlantVillage.
- Utilized a two-stage pipeline for automated QA generation with expert linguistic review.
- Images covered 38 crop species and 89 disease conditions, with questions categorized by complexity and type.
- Dataset underwent iterative review by domain experts for accuracy and relevance.
Main Results:
- Current state-of-the-art vision-language models, including multimodal LLMs, exhibit poor performance on PlantExpertVQA.
- Parameter-efficient fine-tuning of a 2B-parameter model on a subset of the dataset led to significant performance gains.
- Improvements were observed across all question categories, validating the dataset's utility for domain adaptation.
Conclusions:
- PlantExpertVQA is a valuable resource for advancing AI in agriculture, particularly for diagnostic applications.
- Effective domain adaptation is achievable with compact models and targeted fine-tuning on specialized datasets.
- The findings highlight the potential for AI to support farmers in complex decision-making processes.
More Related Videos
Related Concept Videos
Vision
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
Light Acquisition
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
Photoreceptors and Plant Responses to Light
Light plays a significant role in regulating the growth and development of plants. In addition to providing energy for photosynthesis, light provides other important cues to regulate a range of developmental and physiological responses in plants.

