Related Experiment Video
Updated: Jul 1, 2026

07:57
Fish Sperm Assessment Using Software and Cooling Devices
Published on: July 28, 2018
8.7K
Testing the generalizability and effectiveness of deep learning models among clinics: sperm detection as a pilot
Jiaqi Wang1, Yufei Jin1, Aojun Jiang2
1School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China.
Reproductive Biology and Endocrinology : RB&E
|May 22, 2024
Summary
Training deep learning models with diverse imaging data improves their generalizability for sperm analysis in clinical in vitro fertilization (IVF). Richer datasets enhance model precision and recall, ensuring consistent performance across different clinics and hardware.
Area of Science:
- Reproductive Medicine
- Artificial Intelligence in Healthcare
- Computer Vision
Background:
- Deep learning (DL) models are increasingly used to assist in clinical in vitro fertilization (IVF).
- Accurate visual detection of sperm, oocytes, and embryos is a critical first step.
- Variations in clinic-specific hardware and protocols raise concerns about DL model generalizability.
Purpose of the Study:
- To investigate the impact of imaging factors on the generalizability of object detection models.
- To pilot this investigation using sperm analysis as a model system.
- To determine how to improve the robustness of DL models for clinical IVF applications.
Main Methods:
- Ablation studies were conducted on state-of-the-art human sperm detection models.
- Assessed the effect of imaging magnification, mode, and sample preprocessing on model precision and recall.
- Enriched training datasets with diverse imaging conditions and validated performance through internal and multi-center clinical tests.
Main Results:
- Removing specific data subsets (e.g., raw images, 20x magnification) significantly reduced model precision and recall.
- A richly trained model achieved high intraclass correlation coefficients (ICCs) for precision (0.97) and recall (0.97).
- Multi-center validation demonstrated no significant performance differences across clinics, confirming enhanced generalizability.
Conclusions:
- Dataset richness is a critical determinant of DL model generalizability in clinical settings.
- Diverse training data is essential for reliable model evaluation.
- Future DL models in andrology and reproductive medicine require comprehensive feature sets for cross-clinic applicability.

