Related Experiment Video
Updated: Apr 4, 2026

A Semi-high-throughput Imaging Method and Data Visualization Toolkit to Analyze C. elegans Embryonic Development
Published on: October 29, 2019
Investigating discrepancies in accuracy, agreement and interpretability for single-frame embryo classification tasks
Radhika Kakulavarapu1,2, Erwan Delbarre1, Akriti Sharma3
1Department of Life Sciences and Health, Faculty of Health Sciences, OsloMet - Oslo Metropolitan University, Oslo, Norway.
Human embryologists outperformed artificial intelligence models in embryo stage classification accuracy. Comprehensive evaluation frameworks are crucial for the safe development of AI in assisted reproduction technology.
Area of Science:
- Reproductive Medicine
- Medical Imaging
- Artificial Intelligence
Background:
- Artificial intelligence (AI) shows potential in clinical decision support, but requires rigorous evaluation beyond accuracy.
- Assessing AI in assisted reproduction technology necessitates evaluating model agreement with experts and the interpretability of decisions.
- Explainable AI (XAI) methods are vital for understanding AI model behavior in clinical contexts.
Purpose of the Study:
- To evaluate the accuracy and agreement of human embryologists and deep learning (DL) models in classifying embryo developmental stages.
- To explore the interpretability of DL models using explainable AI techniques.
Main Methods:
- A retrospective study analyzed 245 single-frame embryo images.
- Embryo images were classified by three human embryologists and two DL models (ResNet-34, VGG16).
- Evaluation included accuracy, inter-operator agreement (using Cohen's kappa), and interpretability assessment of DL model explanations (Grad-CAMs).
Main Results:
- Human embryologists achieved higher accuracy (89.9%) compared to ResNet-34 (78.8%) and VGG16 (74.3%).
- Overall agreement with the reference standard was excellent for all operators (κ≥0.932), but stage-wise agreement was stronger among embryologists.
- While ResNet-34's explanations were rated more biologically relevant, interpretability did not consistently correlate with accuracy, with weak spatial overlap in explanations, especially at the blastocyst stage.
Conclusions:
- Current AI models demonstrate limitations in accuracy and interpretability for embryo stage classification compared to human experts.
- The study underscores the necessity of integrated evaluation frameworks encompassing accuracy, agreement, and interpretability for AI in assisted reproduction.
- Developing safe and transparent AI tools for assisted reproduction requires robust validation beyond simple accuracy metrics.
More Related Videos
08:33Author Spotlight: Strategies for Mounting Zebrafish Embryos for High-Resolution Multiview Light-Sheet Microscopy — Techniques for Imaging and Image Reconstruction
Published on: July 19, 2024
11:25Quantitative Analysis of Protein Expression to Study Lineage Specification in Mouse Preimplantation Embryos
Published on: February 22, 2016