Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Corrigendum to "4D Flow cardiovascular magnetic resonance consensus statement: 2023 update" [Journal of Cardiovascular Magnetic Resonance 25 (2023) 40].

Journal of cardiovascular magnetic resonance : official journal of the Society for Cardiovascular Magnetic Resonance·2026
Same author

AUTOMATED MEASUREMENTS OF MITRAL AND TRICUSPID ANNULAR DIMENSIONS IN CARDIOVASCULAR MAGNETIC RESONANCE.

Proceedings. IEEE International Symposium on Biomedical Imaging·2026
Same author

TVnet: Automated Time-Resolved Tracking of the Tricuspid Valve Plane in MRI Long-Axis Cine Images with a Dual-Stage Deep Learning Pipeline.

Medical image computing and computer-assisted intervention : MICCAI ... International Conference on Medical Image Computing and Computer-Assisted Intervention·2026
Same author

Atrioventricular area difference assessed by exercise cardiovascular magnetic resonance shows impaired diastolic filling in patients with heart failure.

Journal of applied physiology (Bethesda, Md. : 1985)·2026
Same author

Retrospectively synchronized time-resolved ventricular cine images from 2D real-time exercise cardiac magnetic resonance imaging.

Clinical physiology and functional imaging·2025
Same author

Neural networks with personalized training for improved MOLLI T<sub>1</sub> mapping.

BMC medical imaging·2025

Related Experiment Video

Updated: Jun 18, 2025

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.7K

Random effects during training: Implications for deep learning-based medical image segmentation.

Julius Åkesson1, Johannes Töger1, Einar Heiberg2

  • 1Clinical Physiology, Department of Clinical Sciences Lund, Lund University, Lund, Sweden; Department of Biomedical Engineering, Faculty of Engineering, Lund University, Lund, Sweden.

Computers in Biology and Medicine
|August 3, 2024
PubMed
Summary

Random variations during training create significant performance differences in deep learning segmentation models. Statistical significance is an unreliable indicator of true performance differences between models from the same algorithm.

Keywords:
Deep learningMedical image segmentationPerformance comparisonsRandom seedsRandomness

More Related Videos

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

380
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

2.4K

Related Experiment Videos

Last Updated: Jun 18, 2025

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.7K
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

380
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
05:56

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application

Published on: April 14, 2023

2.4K

Area of Science:

  • Medical image analysis
  • Deep learning
  • Computational neuroscience

Background:

  • Deep learning models for image segmentation exhibit performance variability due to random training effects.
  • This variability can impact the reliability of standard model comparison methods.

Purpose of the Study:

  • To assess the impact of random training effects on the reliability of comparing deep learning segmentation models.
  • To evaluate the effectiveness of common statistical methods in detecting true performance differences.

Main Methods:

  • Utilized nnU-Net with 50 random seeds across three 3D medical image segmentation tasks (brain tumor, hippocampus, cardiac).
  • Assessed segmentation performance using hold-out validation and 5-fold cross-validation.
  • Employed Paired t-test and Wilcoxon signed rank test on Dice scores to measure statistical significance.

Main Results:

  • The top-performing seed significantly outperformed 0-76% of other seeds with hold-out validation.
  • With 5-fold cross-validation, the top seed outperformed 10-38% of other seeds.
  • High rates of statistically significant differences were observed between models trained with the same algorithm.

Conclusions:

  • Random training effects can lead to frequent statistically significant performance differences.
  • Statistical significance is a weak and unreliable indicator of genuine performance superiority between models from the same deep learning algorithm.