Related Experiment Video
Updated: Aug 8, 2026

Fluorimetric Techniques for the Assessment of Sperm Membranes
Published on: November 28, 2018
Evaluation of vision transformers and vision foundation models for sperm morphology analysis
Rawan AlSaad1, Shima Albasha2, Hasan Burjaq2
1AI Center for Precision Health, Weill Cornell Medicine-Qatar, Doha, Qatar.
Background:
Sperm morphology assessment is an essential but highly variable component of semen analysis, influenced by staining, image quality, acquisition conditions, and observer interpretation. Modern pretrained visual encoders may improve the objectivity and reproducibility of this assessment, but their performance and robustness in sperm morphology analysis remain insufficiently characterized.
Objectives:
To compare vision transformers and vision foundation models with established pretrained CNNs for sperm morphology classification; assess the effect of adaptation strategy; and evaluate performance consistency across image sources, generalization to unseen datasets, and the morphological relevance of model attributions.
Methods:
We evaluated binary normal-versus-abnormal sperm morphology classification using 7,770 microscopy images (1,253 normal and 6,517 abnormal sperm images) obtained from three independent institutional sources and representing abnormalities of the sperm head, neck/midpiece, and tail. Ten pretrained vision transformers and vision foundation-models spanning supervised, self-supervised, masked-image, and vision-language pretraining were evaluated using linear probing, light fine-tuning, and full fine-tuning. Five pretrained CNNs served as reference baselines. Primary analyses used source-stratified training, validation, and test partitions within each dataset to assess pooled performance, source-specific performance, and class-specific recall. Secondary analyses evaluated generalization to completely unseen datasets through leave-one-source-out validation and examined model attribution patterns using architecture-appropriate explainability methods.
Results:
SigLIP2 with full fine-tuning and ConvNeXt-Tiny achieved comparable weighted F1 scores (0.958 and 0.955, respectively), while BEiT with full fine-tuning and SigLIP2 with light fine-tuning achieved the highest ROC-AUC (0.982). SigLIP2 with full fine-tuning also produced the highest equal-weight cross-source macro-F1 (0.91) and normal-class recall (0.91). External generalization varied markedly across held-out image sources, with DINOv2 under full fine-tuning achieving ROC-AUCs ranging from 0.654 to 0.928 across the three external evaluations. Explainability analyses generally localized predictions to sperm-related structures.
Conclusions:
Pretraining strategy, adaptation depth, and dataset shift substantially influenced performance. Vision transformers and vision foundation models demonstrated strong potential for AI-assisted sperm morphology assessment, supporting their further development as scalable tools to improve the objectivity, consistency, and efficiency of semen analysis.

