Related Experiment Video
Updated: May 7, 2026

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
Published on: January 13, 2023
Scaling down annotation needs: The capacity of self-supervised learning on diatom classification
Mingkun Tan1, Daniel Langenkämper1, Michael Kloster2
1Biodata Mining Group, Faculty of Technology, University of Bielefeld, 33501 Bielefeld, NRW, Germany.
Abstract:
In the field of life sciences, diatoms are essential biomarkers for assessing environmental health. Recent advancements in deep learning have transformed the traditionally laborious process of diatom classification through light microscopy. However, commonly used supervised learning methodologies necessitate annotated data, demanding the expertise of seasoned professionals. This study introduces self-supervised learning to tackle the challenge of scarce annotation in diatom classification. First, our results reveal that self-supervised pre-trained models considerably enhance the utilization effectiveness of available annotated data, with benefits increasing as the dataset size decreases. Second, fine-tuning our models with a very small labeled dataset (e.g., 50 samples per class) yields macro-average accuracy comparable to full-supervised levels, thereby reducing the reliance on taxonomic experts by approximately . Moreover, extending the pre-training phase to 1600 epochs further reduced the dependency on annotations, achieving comparable accuracy with merely 30 samples per class.
Related Concept Videos
The Scientific Method
Generally, predictions are tested using carefully-designed experiments. Based on the outcome of these...
Key Elements for Plant Nutrition

