Related Experiment Video
Updated: Feb 11, 2026

07:58
Use of Image Cytometry for Quantification of Pathogenic Fungi in Association with Host Cells
Published on: June 19, 2013
13.4K
Naturalgwas: An R package for evaluating genomewide association methods with empirical data
Olivier François1,2,3, Kevin Caye1
1Université Grenoble-Alpes, Grenoble, France.
Molecular Ecology Resources
|April 20, 2018
Summary
This study introduces an R program to simulate phenotypes and estimate the power of genetic association tests, addressing challenges in large-scale studies. The tool helps researchers choose appropriate methods and understand test performance for specific genetic data.
Area of Science:
- Population Genetics
- Statistical Genetics
- Bioinformatics
Background:
- Large-scale genetic association studies face challenges due to geographic variation in genotype frequencies, leading to false positives.
- Existing methods for mitigating these issues lack tools to evaluate the achievable statistical power for specific study designs.
- Population structure and gene-by-environment interactions are significant confounders in genetic association analyses.
Purpose of the Study:
- To present an R program for simulating phenotypes from genotypes to estimate the upper bounds on achievable power for association tests.
- To provide a tool for evaluating the performance of different association testing methods in the presence of population structure and environmental factors.
- To guide researchers in methodological choices for genetic association studies and assess test performance.
Main Methods:
- Development of an R package to simulate phenotypes incorporating population structure and gene-by-environment interactions.
- Implementation of a gold-standard test to evaluate association test power using confounder information.
- Application of the program to simulated data for Arabidopsis thaliana to compare association methods and assess power.
Main Results:
- The developed program successfully simulates phenotypes and estimates power, providing upper bounds for association test performance.
- Simulated data analysis showed that new association methods performed comparably to the gold-standard test in removing confounding factors.
- Statistical power to detect causal variants was influenced by the number of variants and their effect sizes, with specific thresholds identified.
Conclusions:
- The R program offers valuable guidance for selecting appropriate genetic association testing methodologies.
- The tool provides crucial insights into the performance of association tests within user-specific genetic contexts.
- This approach aids in optimizing study design and interpreting results in large-scale genetic association studies.
Keywords:
environmental datagenomewide association studiesgold-standard methodlatent factor modelsphenotypic trait simulationMore Related Videos
Related Concept Videos
Empirical Method to Interpret Standard Deviation
10.3K
The empirical rule, also known as the three-sigma rule, allows a statistician to interpret the standard deviation in a normally distributed dataset. The rule states that 68% of the data lies within one standard deviation from the mean, 95% lies within two standard deviations from the mean, and 99.7% lies within three standard deviations from the mean. Additionally, this rule is also called the 68-95-99.7 rule.
This rule is used widely in statistics to calculate the proportion of data values...
This rule is used widely in statistics to calculate the proportion of data values...
10.3K
DNA Packaging
113.5K
Overview
113.5K
Chromatin Packaging
22.3K
Each human somatic cell contains 6 billion base-pairs of DNA. Each base-pair is 0.34 nm long, which means that each diploid cell contains a staggering 2 meters of DNA. How is such a long DNA strand packed inside a nucleus measuring only 10 - 20 microns in diameter?
The chromatin
In combination with specialized DNA binding protein called Histones, the DNA double helix forms a compact DNA: protein complex called chromatin. The chromatin itself is further compacted into higher-order...
The chromatin
In combination with specialized DNA binding protein called Histones, the DNA double helix forms a compact DNA: protein complex called chromatin. The chromatin itself is further compacted into higher-order...
22.3K
Chromatin Packaging
19.5K
Each human somatic cell contains 6 billion base pairs of DNA. Each base pair is 0.34 nm long, meaning each diploid cell contains a staggering 2 meters of DNA. This long DNA strand is packed inside a nucleus measuring only 10-20 microns in diameter with the help of specialized DNA-binding proteins called histones. Together they form a compact DNA-protein complex called chromatin. The chromatin is further compacted into higher-order structures. The highest level of compaction is achieved during...
19.5K
Chromatin Packaging
9.9K
9.9K
Statistical Methods for Analyzing Epidemiological Data
989
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
989

