Related Experiment Video
Updated: Jun 11, 2025

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
Published on: October 27, 2023
Synthetic data at scale: a development model to efficiently leverage machine learning in agriculture
Jonathan Klein1, Rebekah Waller2, Sören Pirk3
1Computational Sciences Group, King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia.
Abstract:
The rise of artificial intelligence (AI) and in particular modern machine learning (ML) algorithms during the last decade has been met with great interest in the agricultural industry. While undisputedly powerful, their main drawback remains the need for sufficient and diverse training data. The collection of real datasets and their annotation are the main cost drivers of ML developments, and while promising results on synthetically generated training data have been shown, their generation is not without difficulties on their own. In this paper, we present a development model for the iterative, cost-efficient generation of synthetic training data. Its application is demonstrated by developing a low-cost early disease detector for tomato plants (Solanum lycopersicum) using synthetic training data. A neural classifier is trained by exclusively using synthetic images, whose generation process is iteratively refined to obtain optimal performance. In contrast to other approaches that rely on a human assessment of similarity between real and synthetic data, we instead introduce a structured, quantitative approach. Our evaluation shows superior generalization results when compared to using non-task-specific real training data and a higher cost efficiency of development compared to traditional synthetic training data. We believe that our approach will help to reduce the cost of synthetic data generation in future applications.
Related Concept Videos
Light Acquisition
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Plant Breeding and Biotechnology
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
The Central Dogma
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...

