Related Experiment Video
Updated: Aug 14, 2026

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay (EMSA) and DNA-affinity Precipitation Assay (DAPA)
Published on: August 21, 2016
Data-dredging gene-dose analyses in association studies: biases and their corrections
Wen-Chung Lee1, Hsiao-Yuan Huang
1Graduate Institute of Epidemiology, College of Public Health, National Taiwan University, No. 1 Jen-Ai Road, 1st Section, Taipei, Taiwan. wenchung@ha.mc.ntu.edu.tw
Gene-dose analyses in disease risk studies can be biased if "high-risk" genotypes are defined by case-control frequency differences. This study demonstrates that such data-dredging leads to false positives, but a permutation correction method effectively resolves these biases.
Area of Science:
- Genetics
- Biostatistics
- Epidemiology
Background:
- Case-control association studies frequently employ gene-dose analyses to assess the combined effect of multiple genetic loci on disease risk.
- A common methodological pitfall involves defining high-risk genotypes or alleles based on their higher frequency within the case group compared to the control group.
Purpose of the Study:
- To investigate the potential for bias introduced by a specific definition of "high-risk" genotypes in gene-dose analyses.
- To evaluate the impact of this "data-dredging" approach on the accuracy of genetic association studies.
- To propose and validate a statistical method for correcting identified biases.
Main Methods:
- Utilized Monte Carlo simulations to model the "data-dredging" gene-dose analysis under various scenarios, including null hypotheses where no true association exists.
- Developed and implemented a permutation correction method designed to mitigate the identified biases.
- Compared the results of the biased analysis with the corrected analysis.
Main Results:
- Monte Carlo simulations confirmed that the "data-dredging" approach can produce significantly biased results, leading to an overestimation of risk even without a true genetic association.
- The proposed permutation correction method demonstrated a high degree of effectiveness in correcting these biases.
- The corrected analyses yielded more accurate estimates of the joint effect of multiple loci on disease risk.
Conclusions:
- The definition of "high-risk" genotypes based on case-control frequency differences in gene-dose analyses is inherently flawed and can lead to spurious findings.
- A permutation correction method offers a robust solution for rectifying biases in gene-dose analyses.
- Accurate genetic association studies require careful methodological considerations to avoid "data-dredging" and ensure reliable results.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Incomplete Dominance
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...