Related Experiment Video
Updated: Oct 9, 2026

Development of automated imaging and analysis for zebrafish chemical screens.
Published on: June 24, 2010
A self-supervised foundation model for large-scale zebrafish image analysis
Ayse B Oktay1, Arpit Tandon1, Brian E Howard1
1Sciome LLC, Research Triangle Park, NC 27709, USA.
Abstract:
High-throughput zebrafish screening generates large volumes of imaging data, creating a need for scalable computational methods that reduce dependence on manual annotation and repeated task-specific model development. In this study, we develop a self-supervised foundation model for zebrafish image analysis pretrained on more than 468,000 images from seven heterogeneous zebrafish imaging datasets spanning different developmental stages, imaging platforms, modalities, and experimental conditions. By learning a shared representation from diverse unlabeled zebrafish images, the proposed framework provides a reusable encoder that can be adapted to multiple biological endpoints and annotation types, reducing the need to learn independent representations for each downstream task. We evaluated two self-supervised pretraining strategies: Self-Distillation with No Labels (DINO) and a sequential Masked Autoencoder (MAE) followed by DINO (MAE+DINO), across morphology classification, view classification, developmental stage prediction, fish localization, embryo clustering, and anatomical segmentation. The learned representations captured visual similarity, developmental progression, embryo-specific trajectories, phenotypic information, and anatomical structure across heterogeneous imaging data. MAE+DINO achieved the strongest overall performance for morphology classification, view classification, developmental stage prediction, and fine-tuned localization, while DINO provided the strongest frozen representation for localization. For anatomical segmentation, zebrafish specific pretraining improved performance when the encoder was frozen, whereas MAE+DINO and ImageNet initialization achieved comparable overall performance after end-to-end fine-tuning with a SegFormer-style decoder. These results demonstrate that organism-centered self-supervised pretraining can provide reusable and biologically informative visual representations that support diverse zebrafish image-analysis tasks and offer a scalable, annotation-efficient alternative to repeatedly developing task-specific representations.

