Related Experiment Video
Updated: Oct 29, 2025

09:34
Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
34.2K
PIP-SNP: a pipeline for processing SNP data featured as linkage disequilibrium bin mapping, genotype imputing and
Wenchao Zhang1, Yun Kang1, Xinbin Dai1
1Noble Research Institute LLC, 2510 Sam Noble Parkway, Ardmore, OK 73401, USA.
NAR Genomics and Bioinformatics
|July 8, 2021
Summary
This study introduces new methods for analyzing high-dimensional single-nucleotide polymorphism (SNP) data, addressing challenges in genotype imputation and dimension reduction. The developed PIP-SNP pipeline effectively utilizes linkage disequilibrium (LD) for improved SNP data processing.
Area of Science:
- Genetics
- Bioinformatics
- Computational Biology
Background:
- Genome-wide association studies (GWAS) grapple with high-dimensional single-nucleotide polymorphism (SNP) genotype data and missing value imputation.
- Linkage disequilibrium (LD), the correlation between nearby SNPs, presents opportunities for dimension reduction and genotype inference.
Purpose of the Study:
- To develop novel methods for SNP dimension reduction and missing genotype imputation in GWAS.
- To create a streamlined pipeline for processing complex SNP data.
Main Methods:
- Utilized a stochastic process to model SNP signals and proposed autocorrelation measures for SNP information redundancy.
- Constructed LD bins based on autocorrelation coefficients and employed k-nearest neighbors (kNN) for genotype imputation.
- Developed methods for identifying optimal synthetic markers and evaluating information loss during dimension reduction.
Main Results:
- Demonstrated satisfactory performance of the proposed methods on real-world SNP data from rice populations (RIL and HapMap).
- Successfully implemented functional modules into a C/C++ web-based pipeline named PIP-SNP.
- The methods effectively reduce SNP dimensionality while conserving essential information.
Conclusions:
- The developed methods and PIP-SNP pipeline offer efficient solutions for common challenges in GWAS data analysis.
- The approach effectively leverages LD for improved SNP data management and analysis.
- PIP-SNP provides a valuable tool for researchers working with large-scale SNP datasets.
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
17.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.1K
Genome-wide Association Studies-GWAS
14.8K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.8K

