Related Experiment Video
Updated: Sep 9, 2025

Targeted DNA Methylation Analysis by Next-generation Sequencing
Published on: February 24, 2015
Finding easy regions for short-read variant calling from pangenome data
Heng Li1,2,3
1Department of Biomedical Informatics, Harvard Medical School, Boston, MA 02215, USA.
Background:
While benchmarks on short-read variant calling suggest a low error rate below 0.5%, they are only applicable to predefined confident regions. For a human sample without such regions, the error rate could be 10 times higher. Although multiple sets of easy regions have been identified to alleviate the issue, they fail to consider nonreference samples or are biased toward existing short-read data or aligners.
Results:
Here, using hundreds of high-quality human assemblies, we derived a set of sample-agnostic easy regions where short-read variant calling reaches high accuracy. These regions cover 88.2% of GRCh38, 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. They achieve a good balance between coverage and easiness and can be generated for other human assemblies or species with multiple well-assembled genomes.
Conclusions:
This resource provides a convenient and powerful way to filter spurious variant calls for clinical or research human samples.
More Related Videos
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
09:10A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
Published on: May 22, 2018
Related Concept Videos
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Genome Annotation and Assembly