Related Experiment Video
Updated: Jul 19, 2025

08:47
Author Spotlight: UAV Remote Sensing for Efficient Invasive Plant Biomass Estimation
Published on: February 9, 2024
1.5K
Identification of herbarium specimen sheet components from high-resolution images using deep learning.
Karen M Thompson1, Robert Turnbull1, Emily Fitzgerald1
1University of Melbourne Melbourne Victoria Australia.
Ecology and Evolution
|August 17, 2023
Summary
Computer vision models can now extract biodiversity data from herbarium images. A trained YOLOv5 model accurately identifies specimen components, improving digitisation efficiency.
Area of Science:
- Biodiversity informatics
- Computer vision
- Digital specimen imaging
Background:
- Herbarium specimens are crucial biodiversity data sources.
- Manual data extraction from digital images is time-consuming and error-prone.
- Advanced computer vision can automate and enhance data capture from herbarium images.
Purpose of the Study:
- To develop and evaluate an object detection model for identifying components on herbarium specimen sheets.
- To assess the model's generalisability across different herbaria.
- To provide a readily usable model for digitisation efforts.
Main Methods:
- Object detection model (YOLOv5) trained on 3371 annotated images from the University of Melbourne Herbarium (MELU).
- Model validation on 1000 annotated images, with performance metrics including precision, recall, and mean average precision (mAP).
- Generalisation testing on specimens from nine global herbaria and fine-tuning experiments.
Main Results:
- The MELU-trained 'sheet-component' model achieved high performance (mAP0.5-0.95 of 0.847), accurately identifying 11 component types.
- Specific labels like 'institutional' and 'annotation' showed strong prediction accuracy (mAP0.5-0.95 of 0.970 and 0.878).
- The model demonstrated generalisability across diverse herbaria (mAP0.5-0.95 between 0.68-0.89) and could be fine-tuned with minimal data.
Conclusions:
- Object detection models, like the YOLOv5-based 'sheet-component' model, significantly enhance biodiversity data extraction from herbarium images.
- The developed model is accurate, generalisable, and adaptable to new collections with limited retraining.
- Making the trained model weights available can support resource-constrained herbaria in their digitisation workflows.

