Ranking causal variants and associated regions in genome-wide association studies by the support vector machine and
Usman Roshan1, Satish Chikkagoudar, Zhi Wei
1Department of Computer Science, New Jersey Institute of Technology, Newark, NJ, USA. usman@cs.njit.edu
Nucleic Acids Research
|February 15, 2011
Summary
Support vector machine (SVM) and random forest (RF) methods improve causal variant identification when applied to top SNPs. These machine learning approaches enhance association region rankings and prediction power, particularly at stricter significance thresholds.
Area of Science:
- Genetics and Bioinformatics
- Statistical Genomics
- Machine Learning in Biology
Background:
- Identifying causal genetic variants is crucial for understanding complex diseases.
- Traditional methods like chi-squared statistics rely on single SNP associations.
- Machine learning offers potential for improved variant and region identification.
Purpose of the Study:
- To compare the performance of chi-squared statistics, support vector machine (SVM), and random forest (RF) in identifying causal variants and associated regions.
- To evaluate the impact of different SNP selection cutoffs on the performance of SVM and RF.
- To assess the utility of these methods for real-world applications, including type 1 diabetes genetics and disease risk prediction.
Main Methods:
- Utilized simulated and real genetic datasets.
- Ranked single nucleotide polymorphisms (SNPs) using 1 df chi-squared statistic, SVM, and RF.
- Applied SVM and RF to top SNPs selected based on chi-squared ranking and Bonferroni correction (2r, 5r, 10r cutoffs).
Main Results:
- SVM and RF applied to top 2r chi-squared ranked SNPs improved causal variant and region rankings and increased power in simulated data.
- Performance and stability of SVM and RF rankings decreased with higher cutoffs (5r, 10r).
- Comparison of previously replicated SNPs in type 1 diabetes data showed differences in rankings across methods.
Conclusions:
- SVM and RF demonstrate superior performance over traditional chi-squared statistics for identifying causal variants and associated regions, especially with optimized SNP selection.
- The effectiveness of machine learning methods is sensitive to the stringency of SNP selection criteria.
- The developed methods show promise for genetic association studies and disease risk prediction, with available software and a webserver.
Related Concept Videos
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Principles of Pharmacogenetics: Types of Genetic Variants
The human genome is over 99.9% identical between individuals, yet genetic differences exist at millions of bases. The human genome contains approximately 3 million variant positions per individual, many of which are heterozygous, contributing to genetic diversity and individual traits. Genetic variations include single-nucleotide polymorphisms (SNPs), insertions, deletions, and copy number variations (CNVs).SNPs, the most common variation, involve single-base changes in DNA. These can be...
Single Nucleotide Polymorphisms-SNPs
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
Genetic Variation
Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles, which...
Genes exist in different versions called alleles, which...

