Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

290
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
290
Sensitivity, Specificity, and Predicted Value01:13

Sensitivity, Specificity, and Predicted Value

664
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
664

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Efficient Deep Learning Models for Predicting Individualized Task Activation From Resting-State Functional Connectivity.

Human brain mapping·2026
Same author

Leveraging Machine Learning to Advance Alcohol Research: Current Applications, Challenges, and Opportunities.

Alcohol research : current reviews·2026
Same author

A view-engage-predict framework for enhancing brain-behavior mapping with naturalistic movie-watching fMRI.

Communications biology·2026
Same author

A framework of digital biomarkers for neurodegenerative diseases.

Nature reviews bioengineering·2026
Same author

SocialGen: Modeling Multi-Human Social Interaction with Language Models.

Proceedings. International Conference on 3D Vision·2026
Same author

Anxiety Symptoms in Preschool Children Born Very Preterm: Associations with Cognition and Neonatal Striatal Volumes.

Children (Basel, Switzerland)·2026

Related Experiment Video

Updated: Sep 12, 2025

Basics of Multivariate Analysis in Neuroimaging Data
06:35

Basics of Multivariate Analysis in Neuroimaging Data

Published on: July 24, 2010

17.0K

Statistical variability in comparing accuracy of neuroimaging based classification models via cross validation.

Bahram Jafrasteh1, Ehsan Adeli2,3, Kilian M Pohl2

  • 1Department of Radiology, Weill Cornell Medicine, New York, NY, USA. baj4003@med.cornell.edu.

Scientific Reports
|August 6, 2025
PubMed
Summary

Comparing machine learning (ML) model accuracy in biomedical research is challenging, especially with cross-validation (CV). This study proposes a framework to assess CV impact on statistical significance, highlighting variability and the need for rigorous comparison practices.

Keywords:
Cross validationMachine learningReproducibility crisisStatistical hypothesis testing

More Related Videos

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

15.8K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.9K

Related Experiment Videos

Last Updated: Sep 12, 2025

Basics of Multivariate Analysis in Neuroimaging Data
06:35

Basics of Multivariate Analysis in Neuroimaging Data

Published on: July 24, 2010

17.0K
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
14:27

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data

Published on: June 26, 2013

15.8K
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.9K

Area of Science:

  • Biomedical research
  • Machine learning applications
  • Neuroimaging analysis

Background:

  • Machine learning (ML) advances classification accuracy in biomedical research.
  • Rigorous comparison of ML model accuracy is essential but challenging.
  • Cross-validation (CV) introduces complexities in statistical significance testing for ML models.

Purpose of the Study:

  • To highlight practical challenges in quantifying statistical significance of accuracy differences between ML models using CV.
  • To propose an unbiased framework for assessing the impact of CV setups on statistical significance.
  • To address the need for more rigorous practices in biomedical ML research to mitigate reproducibility issues.

Main Methods:

  • Developed an unbiased framework to evaluate the influence of CV configurations on statistical significance.
  • Applied the framework to three public neuroimaging datasets.
  • Analyzed the impact of data properties, testing procedures, and CV choices on detecting significant differences.

Main Results:

  • Demonstrated substantial variability in detecting significant ML model differences based on data properties, testing procedures, and CV configurations.
  • Re-emphasized known flaws in current p-value computations for comparing model accuracies.
  • Showed that factors influencing significance are often overlooked in ML-based biomedical studies.

Conclusions:

  • Variability in significance testing can lead to p-hacking and inconsistent conclusions on model improvement.
  • More rigorous practices are urgently needed for comparing ML models in biomedical research.
  • Implementing robust comparison methods is crucial for addressing the reproducibility crisis in the field.