Detecting disease gene in DNA haplotype sequences by nonparametric dissimilarity test
Ao Yuan1, Qingqi Yue, Victor Apprey
1National Human Genome Center, Department of Community Health and Family Medicine, Howard University, Washington, DC, USA. ayuan@howard.edu
Human Genetics
|June 30, 2006
Summary
This study introduces a novel weighted U-statistic for haplotype-based association studies. The new method enhances detection power for complex diseases by comparing genetic dissimilarity, overcoming limitations of existing U-statistic tests.
Area of Science:
- Genetics and Bioinformatics
- Statistical Genetics
- Complex Disease Association Studies
Background:
- Haplotype-based association studies are crucial for understanding complex diseases.
- Existing U-statistic methods for comparing genetic similarity have detection limitations.
- A need exists for more powerful and robust statistical tests in genetic association.
Purpose of the Study:
- To propose a novel weighted U-statistic for haplotype-based association analysis.
- To enhance the power and eliminate blind spots in detecting genetic associations.
- To directly compare genetic dissimilarity between case and control populations.
Main Methods:
- Developed a new weighted U-statistic formulation.
- Analyzed the asymptotic properties of the test statistic under null and alternative hypotheses.
- Conducted simulation studies to evaluate performance against existing methods.
Main Results:
- The proposed weighted U-statistic is asymptotically a linear combination of absolute normal variables.
- The test statistic shows a strict rightward shift under the alternative hypothesis, indicating no blind detection areas.
- Simulation results demonstrate superior robustness and power compared to existing U-statistic variants.
Conclusions:
- The novel weighted U-statistic effectively addresses limitations of current methods in complex disease association studies.
- This approach offers a more powerful and comprehensive tool for analyzing haplotype data.
- The method provides a robust and sensitive means for identifying genetic associations.
Related Concept Videos
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Single Nucleotide Polymorphisms-SNPs
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Test for Homogeneity
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can be stated as...

