Related Experiment Video
Updated: Jul 13, 2026

A Novel Bayesian Change-point Algorithm for Genome-wide Analysis of Diverse ChIPseq Data Types
Published on: December 10, 2012
Haplotype inference using a Bayesian Hidden Markov model
Shuying Sun1, Celia M T Greenwood, Radford M Neal
1Mathematical Biosciences Institute, The Ohio State University, Columbus, Ohio 43210, USA. ssun@mbi.osu.edu
Abstract:
Knowledge of haplotypes is useful for understanding block structure in the genome and disease risk associations. Direct measurement of haplotypes in the absence of family data is presently impractical, and hence, several methods have been developed for reconstructing haplotypes from population data. We have developed a new population-based method using a Bayesian Hidden Markov model for the source of the ancestral haplotype segments. In our Bayesian model, a higher order Markov model is used as the prior for ancestral haplotypes, to account for linkage disequilibrium. Our model includes parameters for the genotyping error rate, the mutation rate, and the recombination rate at each position. Computation is done by Markov Chain Monte Carlo using the forward-backward algorithm to efficiently sum over all possible state sequences of the Hidden Markov model. We have used the model to reconstruct the haplotypes of 129 children at a region on chromosome 5 in the data set of Daly et al. [2001] (for which true haplotypes are obtained based on parental genotypes) and of 30 children at selected regions in the CEU and YRI data of the HAPMAP project. The results are quite close to the family-based reconstructions and comparable with the state-of-the-art PHASE program. Our haplotype reconstruction method does not require division of the markers into small blocks of loci. The recombination rates inferred from our model can help to predict haplotype block boundaries, and estimate recombination hotspots.
Related Concept Videos
Genetic Drift
Mutation, Gene Flow, and Genetic Drift
Hardy-Weinberg Principle
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Statistical Hypothesis Testing
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Evolutionary Relationships through Genome Comparisons

