Related Experiment Video
Updated: Jun 26, 2026

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay (EMSA) and DNA-affinity Precipitation Assay (DAPA)
Published on: August 21, 2016
Variable locus length in the human genome leads to ascertainment bias in functional inference for non-coding elements
1Computational Biology Branch, National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, 8600 Rockville Pike, Bethesda, MD 20894, USA.
Functional inference for non-coding DNA is biased by gene annotation databases. We developed correction coefficients to account for non-coding DNA length variability, eliminating ascertainment bias for accurate functional characterization.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Functional gene annotation databases infer biological function by identifying over- and underrepresented attributes.
- This approach is problematic for non-coding DNA due to variable sequence lengths, leading to biased functional conclusions based on neighboring genes.
Purpose of the Study:
- To assess systematic bias in Gene Ontology (GO) categories when inferring function for non-coding elements based on nearest gene annotations.
- To develop a correction method to eliminate ascertainment bias in non-coding DNA functional inference.
Main Methods:
- Systematic bias in GO categories was assessed using the hypergeometric test.
- Non-coding elements were randomly sampled from the human genome, and their function was inferred from the closest genes.
- Correction coefficients were developed to adjust GO category probabilities based on non-coding DNA length variability.
Main Results:
- Certain GO categories ('cell adhesion', 'nervous system development', 'transcription factor activities') were systematically overrepresented, while others ('olfactory receptor activity') were underrepresented in non-coding elements.
- A novel set of correction coefficients was introduced to account for non-coding DNA length variability.
- The proposed correction effectively eliminated ascertainment bias in functional characterization.
Conclusions:
- Functional inference for non-coding elements using standard gene annotation databases requires specific corrections.
- The developed correction coefficients accurately adjust for ascertainment bias, enabling more reliable functional characterization of non-coding DNA.
- This approach is generalizable to other gene annotation databases.
More Related Videos
Related Concept Videos
Non-LTR Retrotransposons
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
lncRNA - Long Non-coding RNAs
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Single Nucleotide Polymorphisms-SNPs
Genome Size and the Evolution of New Genes

