Related Experiment Video
Updated: Aug 7, 2026

A Strategy to Identify de Novo Mutations in Common Disorders such as Autism and Schizophrenia
Published on: June 15, 2011
A model-based approach to selection of tag SNPs
Pierre Nicolas1, Fengzhu Sun, Lei M Li
1Molecular and Computational Biology Program, Department of Biological Sciences, University of Southern California, Los Angeles, USA. pierre.nicolas@jouy.inra.fr
Selecting tag Single Nucleotide Polymorphisms (SNPs) using the Li and Stephens hidden Markov model maximizes information content, outperforming existing methods for genetic association studies. This approach improves tag set quality assessment and SNP prediction.
Area of Science:
- Genomics
- Bioinformatics
- Statistical Genetics
Background:
- Single Nucleotide Polymorphisms (SNPs) are common variations in the human genome.
- Identifying tag SNPs is crucial for efficient genetic association studies by capturing haplotype information.
- Optimal tag SNP selection requires robust probabilistic models that accurately represent Linkage Disequilibrium (LD).
Purpose of the Study:
- To develop and evaluate tag SNP selection methods based on entropy maximization and probabilistic models.
- To compare the performance of different models, including the Li and Stephens hidden Markov model, for tag SNP selection.
- To assess the informativeness and predictive ability of selected tag SNP sets.
Main Methods:
- Computed description code-lengths of SNP data using various probabilistic models.
- Developed tag SNP selection strategies based on entropy maximization.
- Utilized datasets from the HapMap and ENCODE projects for model evaluation.
Main Results:
- The Li and Stephens hidden Markov model demonstrated superior performance in describing SNP data and maximizing information content.
- Tag SNP sets selected using this model showed enhanced ability to predict tagged SNPs compared to other methods.
- Information content, evaluated with a good model, proved more sensitive for assessing tag set quality than prediction rates.
Conclusions:
- Tag sets selected via the Li and Stephens model significantly outperform those from existing methods.
- Haplotype informativeness, assessed through robust models, is key for effective tag SNP selection, even without direct haplotype genotyping.
- Haplotype phase uncertainty has minimal impact on the predictive power of well-selected tag SNP sets.
More Related Videos
05:53Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
09:33Genetic Profiling and Genome-Scale Dropout Screening to Identify Therapeutic Targets in Mouse Models of Malignant Peripheral Nerve Sheath Tumor
Published on: August 25, 2023
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
Tagging and Fusion Proteins
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Pharmacogenomics: Identification of New Drug Targets
Modern Molecular Taxonomy