Related Experiment Video
Updated: Nov 9, 2025

10:25
Deep Learning-Based Segmentation of Cryo-Electron Tomograms
Published on: November 11, 2022
9.9K
CEM500K, a large-scale heterogeneous unlabeled cellular electron microscopy image dataset for deep learning
Ryan Conrad1,2, Kedar Narayan1,2
1Center for Molecular Microscopy, Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda, United States.
Elife
|April 8, 2021
Summary
A new dataset, CEM500K, enables effective pre-training for deep learning (DL) models in cellular electron microscopy (EM) segmentation. This approach improves model generalization and achieves state-of-the-art results on benchmark tasks.
Area of Science:
- Cellular Biology
- Microscopy
- Artificial Intelligence
Background:
- Automated segmentation of cellular electron microscopy (EM) data is difficult.
- Supervised deep learning (DL) models require extensive region-of-interest (ROI) annotations and do not generalize well to new datasets.
- Unsupervised DL methods need pre-training data, but existing EM datasets are large, homogeneous, computationally expensive to train on, and offer limited value for diverse biological contexts.
Purpose of the Study:
- To develop a more effective pre-training strategy for deep learning (DL) models in cellular electron microscopy (EM) segmentation.
- To create a versatile and computationally efficient dataset for pre-training EM segmentation models.
- To enhance the generalization capabilities and performance of DL models across various EM datasets.
Main Methods:
- Curated a novel dataset, CEM500K, comprising 0.5 million unique 2D cellular EM images from diverse sources.
- Utilized CEM500K for pre-training DL models, focusing on feature learning resilient to image augmentations.
- Evaluated transfer learning performance of CEM500K pre-trained models on multiple benchmark EM segmentation tasks.
Main Results:
- Models pre-trained on CEM500K learned biologically relevant features.
- Pre-trained models demonstrated robustness against image augmentations.
- State-of-the-art results were achieved on six public and one new benchmark EM segmentation tasks using transfer learning from CEM500K.
Conclusions:
- CEM500K provides a computationally efficient and effective resource for pre-training DL models in EM.
- Pre-training on CEM500K significantly improves model generalization and segmentation performance.
- The CEM500K dataset, pre-trained models, and curation pipeline are released to benefit the EM research community.

