Related Experiment Video
Updated: Aug 5, 2025

12:06
Analyzing Mitochondrial Morphology Through Simulation Supervised Learning
Published on: March 3, 2023
4.1K
Impact of Training Data, Ground Truth and Shape Variability in the Deep Learning-Based Semantic Segmentation of HeLa
Cefa Karabağ1, Mauricio Alberto Ortega-Ruíz1,2, Constantino Carlos Reyes-Aldasoro1
1giCentre, Department of Computer Science, School of Science and Technology, City, University of London, London EC1V 0HB, UK.
Journal of Imaging
|March 28, 2023
Summary
Increasing training data for U-Net segmentation of HeLa cells significantly improved accuracy. Automatically generated data, when combined with manual data, yielded the best segmentation results for electron microscopy images.
Area of Science:
- * Biomedical image analysis
- * Deep learning in microscopy
Background:
- * Accurate cell segmentation is crucial for biological research.
- * Deep learning models like U-Net show promise for image segmentation.
- * Evaluating the impact of training data size and quality is essential for model performance.
Purpose of the Study:
- * To investigate how training data amount and variability affect U-Net segmentation of HeLa cells.
- * To assess the correctness of manually generated ground truth data.
- * To compare U-Net performance against traditional image processing algorithms.
Main Methods:
- * Used 3D electron microscopy images of HeLa cells (8192×8192×517).
- * Cropped and manually delineated a region of interest (ROI) for ground truth.
- * Generated data/label pairs for nucleus, nuclear envelope, cell, and background.
- * Trained U-Net architectures with varying numbers of data pairs (36,000 to 270,000).
- * Compared results from manually segmented and automatically generated data.
Main Results:
- * Increased training data pairs led to improved accuracy and Jaccard similarity index for ROI segmentation.
- * U-Net trained with automatically generated data outperformed manual segmentation on larger slices.
- * Combining manually and automatically generated data (270,000 pairs) yielded the best segmentation performance.
Conclusions:
- * The quantity and diversity of training data significantly impact U-Net segmentation performance.
- * Automatically generated training data can effectively represent cell variability for improved segmentation.
- * Optimizing training data generation and quantity is key for robust deep learning-based cell image analysis.

