Risk prediction and marker selection in nonsynonymous single nucleotide polymorphisms using whole genome sequencing
Young-Sup Lee1, KyeongHye Won1, Donghyun Shin2,3
1Department of Animal Biotechnology, Jeonbuk National University, Jeonju, Republic of Korea.
Animal Cells and Systems
|January 18, 2021
Summary
This study introduces a genome-wide approach using nonsynonymous single nucleotide polymorphisms (nsSNPs) to predict individual health risks. Machine learning identified key nsSNPs impacting DNA metabolism, immunity, and reproduction.
Area of Science:
- Genetics
- Bioinformatics
- Genomics
Background:
- Nonsynonymous single nucleotide polymorphisms (nsSNPs) can alter protein structure and function, potentially leading to deleterious effects.
- Genome-wide studies focusing on the impact of nsSNPs remain relatively rare despite their known functional consequences.
Purpose of the Study:
- To develop a method for predicting individual risk scores based on the deleterious effects of nsSNPs.
- To identify an optimal subset of nsSNPs that effectively represents the cumulative nsSNP effect using machine learning.
- To explore the biological functions associated with selected nsSNPs through gene ontology analysis.
Main Methods:
- Prediction of deleterious effects for a large set of nsSNPs.
- Application of machine learning to select a representative nsSNP subset from a larger dataset.
- Gene ontology analysis to identify enriched biological processes among the selected nsSNPs.
Main Results:
- A total of 16,100 nsSNPs were identified as optimal representatives out of 89,519 initially regressed nsSNPs.
- Gene ontology analysis revealed significant enrichment in DNA metabolic processes, chemokine- and immune-related functions, and reproduction.
- A risk score prediction model was developed based on the identified nsSNPs.
Conclusions:
- The developed risk score prediction and nsSNP marker selection methods offer a novel approach for genome-wide association studies.
- These findings are expected to advance breeding science by providing better genetic markers.
- The study highlights the importance of nsSNPs in fundamental biological processes and individual risk assessment.
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
17.4K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.4K
Genome-wide Association Studies-GWAS
14.9K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.9K
Comparing Copy Number Variations and SNPs
18.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.4K


