Related Experiment Video
Updated: Nov 1, 2025

08:47
Author Spotlight: UAV Remote Sensing for Efficient Invasive Plant Biomass Estimation
Published on: February 9, 2024
1.8K
Image-Based Automated Species Identification: Can Virtual Data Augmentation Overcome Problems of Insufficient
Morris Klasen1, Dirk Ahrens2, Jonas Eberle2,3
1Department of Computer Science IV, University of Bonn, Endenicher Allee 19A, 53115 Bonn, Germany.
Systematic Biology
|June 18, 2021
Summary
Data augmentation techniques, including image rotation and generative adversarial networks, improve automated species identification accuracy, especially for rare species with limited data. This approach overcomes challenges in machine learning for biodiversity research.
Area of Science:
- Biodiversity Informatics
- Computational Biology
- Machine Learning Applications
Background:
- Automated species identification is crucial but challenging, particularly for rare species with limited sampling data.
- Scarcity of infraspecific data hinders the performance of machine learning models in distinguishing between closely related species.
- Existing methods struggle with low or exaggerated interspecific morphological variation.
Purpose of the Study:
- To evaluate the effectiveness of a data augmentation strategy for enhancing automated visual species identification.
- To address the challenge of limited infraspecific sampling data in machine learning for species identification.
- To improve the accuracy of automated species identification, especially for rare and under-sampled taxa.
Main Methods:
- Employed a stepwise data augmentation approach including image rotation, visual augmentation, and generative adversarial networks (GANs).
- Extracted descriptive feature vectors from VGG-16 convolutional neural network bottleneck features, reduced dimensionality using Global Average Pooling and Principal Component Analysis (PCA).
- Utilized synthetic oversampling in feature space to augment limited training data.
Main Results:
- The proposed data augmentation approach significantly outperformed a deep learning baseline using non-augmented data.
- The augmentation method also showed superior performance compared to a traditional 2D morphometric approach (Procrustes analysis).
- The study demonstrated improved identification accuracy across diverse datasets including beetles, bees, and butterflies.
Conclusions:
- Data augmentation is an effective strategy to overcome the limitations of scarce training data in automated species identification.
- The combined approach of deep learning, GANs, and synthetic oversampling offers a robust solution for identifying rare species.
- This methodology holds significant promise for advancing biodiversity research and conservation efforts through improved automated identification tools.

