Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Survival Tree01:19

Survival Tree

Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a survival tree begins...
Prediction Intervals01:03

Prediction Intervals

The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
The...
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Real-time volumetric imaging of cells and molecules in deep tissues with Takoyaki ultrasound.

Nature communications·2026
Same author

Association between smartphone overdependence and alcohol and tobacco use behaviors among adolescents in Korea.

Scientific reports·2026
Same author

Statistical BURST imaging for high-fidelity biomolecular ultrasound.

bioRxiv : the preprint server for biology·2026
Same author

Necroptosis induced by MLKL overexpression in liver triggers cellular senescence and leads to chronic inflammation and fibrosis.

GeroScience·2025
Same author

Seizure evolution in a mouse model of West syndrome involves complex and time-dependent synapse remodeling, gliosis and alterations in lipid metabolism.

PLoS biology·2025
Same author

Sono-uncaging for Spatiotemporal Control of Chemical Reactivity.

Journal of the American Chemical Society·2025

Related Experiment Video

Updated: Jul 6, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

Mistakes in validating the accuracy of a prediction classifier in high-dimensional but small-sample microarray data.

Sunho Lee1

  • 1Department of Applied Mathematics, Sejong University, Seoul, South Korea. leesh@sejong.ac.kr

Statistical Methods in Medical Research
|April 1, 2008
PubMed
Summary

Developing accurate gene expression classifiers for clinical use requires careful validation. This study emphasizes assessing classifier reproducibility and avoiding cross-validation pitfalls to ensure reliable diagnostic tools.

More Related Videos

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
03:37

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets

Published on: March 1, 2024

Related Experiment Videos

Last Updated: Jul 6, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
03:37

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets

Published on: March 1, 2024

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Medical Informatics

Background:

  • Gene expression microarray studies aim to create accurate clinical classifiers.
  • Overfitting is a risk with high-dimensional gene data and small sample sizes, leading to non-reproducible results.
  • Ensuring classifier reproducibility is crucial for clinical adoption.

Purpose of the Study:

  • To discuss appropriate methods for validating gene expression classifiers.
  • To estimate the predictive accuracy of developed classifiers.
  • To review common mistakes in cross-validation for classifier development.

Main Methods:

  • Review of validation and accuracy estimation methods for gene expression classifiers.
  • Analysis of published articles in prominent medical journals to identify cross-validation errors.
  • Discussion of strategies to prevent non-reproducible classifier results.

Main Results:

  • Overfitting in gene expression classification can lead to unreliable and non-reproducible outcomes.
  • Inappropriate cross-validation techniques can yield misleading accuracy estimates.
  • Identified common errors in cross-validation processes within published research.

Conclusions:

  • Rigorous validation and reproducibility assessment are essential for clinical gene expression classifiers.
  • Awareness of cross-validation pitfalls is necessary to avoid erroneous results.
  • Preventing indefinite classifier development ensures appropriate patient treatment decisions.