Related Experiment Video
Updated: May 25, 2026

09:33
Genetic Profiling and Genome-Scale Dropout Screening to Identify Therapeutic Targets in Mouse Models of Malignant Peripheral Nerve Sheath Tumor
Published on: August 25, 2023
The illusion of distribution-free small-sample classification in genomics
Edward R Dougherty1, Amin Zollanvari, Ulisses M Braga-Neto
1Department of Electrical and Computer Engineering, Texas A&M University.
Current Genomics
|February 2, 2012
Summary
Accurate error estimation is crucial for bioinformatics classification. Without stated distributional assumptions and accuracy measures, classification in small-sample biology is unreliable.
Area of Science:
- Bioinformatics
- Genomic Data Analysis
- Biostatistics
Background:
- Classification is vital in bioinformatics for distinguishing phenotypes, especially diseases, using genomic data.
- Current methods often lack robust error estimation rules and theoretical underpinnings for accuracy.
- The reliability of a classifier hinges on its error rate, yet this is frequently overlooked.
Purpose of the Study:
- To highlight the critical need for accurate error estimation in bioinformatics classification.
- To underscore the limitations of current methods in small-sample, high-throughput genomic data settings.
- To advocate for the explicit statement of distributional assumptions and error estimation accuracy.
Main Methods:
- Review of existing classification and error estimation methodologies in bioinformatics.
- Analysis of the impact of distributional assumptions on classification accuracy.
- Examination of the challenges posed by small sample sizes in genomic data.
Main Results:
- Many bioinformatics classification studies lack rigorous error estimation and accuracy assessment.
- Cross-validation on the same dataset provides unreliable error estimates without distributional assumptions.
- Distribution-free bounds for error estimation are often impractical for small sample sizes.
Conclusions:
- Scientifically meaningful classification in high-throughput, small-sample biology is illusory without stated distributional assumptions and accuracy measures.
- Accurate error estimation is epistemologically dependent on understanding the underlying data distribution.
- Future bioinformatics research must prioritize robust error estimation and theoretical validation.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Genomics
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...