Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Cluster Sampling Method01:20

Cluster Sampling Method

11.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.0K
One-Way ANOVA: Equal Sample Sizes01:15

One-Way ANOVA: Equal Sample Sizes

3.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.2K
Bootstrapping01:24

Bootstrapping

784
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
784
One-Way ANOVA: Unequal Sample Sizes01:15

One-Way ANOVA: Unequal Sample Sizes

5.8K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.8K
Random Sampling Method01:09

Random Sampling Method

11.8K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
11.8K
Sampling Methods: Overview01:06

Sampling Methods: Overview

3.7K
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling. 
In analytical chemistry, the choice of...
3.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

SNP-based prediction of schizophrenia using machine learning.

Bioinformatics advances·2026
Same author

Predicting Intensive Care Unit Admission in COVID-19-Infected Pregnant Women Using Machine Learning.

Journal of clinical medicine·2025
Same author

Pathway-based analyses of gene expression profiles at low doses of ionizing radiation.

Frontiers in bioinformatics·2024
Same author

Optimal decision-making in high-throughput virtual screening pipelines.

Patterns (New York, N.Y.)·2023
Same author

Knowledge-driven learning, optimization, and experimental design under uncertainty for materials discovery.

Patterns (New York, N.Y.)·2023
Same author

Diagnosis of Endometriosis Based on Comorbidities: A Machine Learning Approach.

Biomedicines·2023

Related Experiment Video

Updated: Apr 25, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.3K

Cross-validation under separate sampling: strong bias and how to correct it.

Ulisses M Braga-Neto1, Amin Zollanvari1, Edward R Dougherty1

  • 1Department of Electrical and Computer Engineering, Center for Bioinformatics and Genomic Systems Engineering and Department of Statistics, Texas A&M University, College Station, TX, 77843, USA Department of Electrical and Computer Engineering, Center for Bioinformatics and Genomic Systems Engineering and Department of Statistics, Texas A&M University, College Station, TX, 77843, USA.

Bioinformatics (Oxford, England)
|August 16, 2014
PubMed
Summary

Classical cross-validation can be biased in bioinformatics when using separate sampling. A new method is proposed to provide almost unbiased error estimation, crucial for accurate pattern recognition results.

More Related Videos

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
08:56

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates

Published on: January 13, 2023

2.5K
Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

21.0K

Related Experiment Videos

Last Updated: Apr 25, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.3K
Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates
08:56

Automatic Image Processing to Determine the Community Size Structure of Riverine Macroinvertebrates

Published on: January 13, 2023

2.5K
Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

21.0K

Area of Science:

  • Bioinformatics
  • Pattern Recognition
  • Statistical Learning Theory

Background:

  • Cross-validation is a standard technique for estimating prediction error in pattern recognition.
  • The assumption of 'almost unbiased' error estimation holds for random sampling but not for separate sampling, common in bioinformatics.

Purpose of the Study:

  • To investigate the bias of classical cross-validation under separate sampling.
  • To develop and validate a new cross-validation error estimator for separate sampling scenarios.

Main Methods:

  • Analytical and numerical methods were used to demonstrate the bias.
  • A novel separate-sampling cross-validation error estimator was proposed and theoretically analyzed.
  • The proposed estimator was proven to satisfy an 'almost unbiased' theorem.

Main Results:

  • Classical cross-validation exhibits significant bias under separate sampling, influenced by sampling ratios and population probabilities.
  • The new separate-sampling cross-validation estimator achieves 'almost unbiased' error estimation.
  • Case studies using existing data revealed drastic result changes when the appropriate cross-validation method was applied.

Conclusions:

  • Separate sampling in bioinformatics requires specialized cross-validation techniques.
  • The proposed estimator offers a reliable alternative to classical cross-validation for separate sampling.
  • Accurate error estimation is critical for robust pattern recognition in bioinformatics.