Related Experiment Video
Updated: Sep 5, 2025

12:39
A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
Published on: December 10, 2012
11.4K
P-smoother: efficient PBWT smoothing of large haplotype panels
William Yue1, Ardalan Naseri1, Victor Wang1
1School of Biomedical Informatics, University of Texas Health Science Center at Houston, Houston, TX 77030, USA.
Bioinformatics Advances
|July 5, 2022
Summary
P-smoother corrects errors in large haplotype panels using a novel PBWT-based algorithm. This improves the identification of identical-by-descent segments and enables state-of-the-art multiway IBD detection for biobank-scale data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Large haplotype panels are crucial for genetic studies.
- Exact haplotype matching is hindered by recent mutations and genotyping errors.
- Existing algorithms passively tolerate mismatches, limiting accuracy.
Purpose of the Study:
- To develop an efficient algorithm for correcting mismatches in large haplotype panels.
- To enhance the accuracy of identifying shared haplotypes and identical-by-descent (IBD) segments.
- To enable state-of-the-art performance in multiway IBD segment detection.
Main Methods:
- Proposed P-smoother, a PBWT-based smoothing algorithm.
- Implemented bidirectional PBWT scanning to correct mismatching alleles using the IBD prior.
- Evaluated performance on simulated data and UK Biobank data.
Main Results:
- P-smoother reliably corrected 85% of errors in a simulated panel with a 0.2% error rate.
- Smoothed panels enabled PBWT algorithms to identify more pairwise IBD segments.
- The PS-cluster algorithm achieved state-of-the-art performance in identifying multiway IBD segments.
Conclusions:
- P-smoother effectively corrects errors in haplotype panels, improving IBD detection.
- PS-cluster demonstrates state-of-the-art efficiency and accuracy for multiway IBD analysis.
- P-smoother enables new possibilities for error-tolerant algorithms in biobank-scale genomics.
Related Concept Videos
Genome-wide Association Studies-GWAS
14.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.1K
Single Nucleotide Polymorphisms-SNPs
15.7K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.7K

