Related Experiment Video
Updated: Jul 26, 2025

Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
Discovery of non-reference processed pseudogenes in the Swedish population
Esmee Ten Berk de Boer1, Kristine Bilgrav Saether1,2, Jesper Eisfeldt1,2,3
1Department of Molecular Medicine and Surgery, Center for Molecular Medicine, Karolinska Institutet, Stockholm, Sweden.
This study reveals thousands of new processed pseudogenes in the human genome, highlighting their variability and potential use in DNA testing. The developed pipeline also improves accuracy in structural variation analysis.
Area of Science:
- Genomics
- Human Genetics
- Bioinformatics
Background:
- The human genome contains a vast majority of non-coding DNA, historically understudied and often termed 'junk DNA'.
- Pseudogenes, non-functional gene copies, are a significant non-coding feature, with processed pseudogenes arising from retrotransposition of mRNA.
- The variability and distribution of processed pseudogenes across human populations remain largely unknown.
Purpose of the Study:
- To develop and apply a custom pipeline for identifying processed pseudogenes in whole genome sequencing data.
- To discover novel processed pseudogenes absent from the current human genome reference (GRCh38).
- To analyze the variability and frequency of processed pseudogenes for potential applications in DNA testing and population genetics.
Main Methods:
- Applied a custom processed pseudogene pipeline to whole genome sequencing data from 3,500 individuals (1000 Genomes dataset and Swedish cohort).
- Analyzed sequence data to identify and position processed pseudogenes, including those not present in the GRCh38 reference.
- Evaluated the impact of processed pseudogenes on common structural variant calling algorithms.
Main Results:
- Discovered over 3,000 processed pseudogenes not annotated in the GRCh38 human genome reference.
- Successfully positioned 74% of detected processed pseudogenes, enabling formation analysis.
- Identified that common structural variant callers misclassify processed pseudogenes as deletions, leading to false positive truncating variant predictions.
- Documented significant variability in non-reference processed pseudogenes across populations, with potential utility as population-specific markers.
Conclusions:
- Processed pseudogenes are highly diverse and actively generated within the human genome.
- The developed pipeline effectively identifies novel processed pseudogenes and can reduce false positives in structural variation analysis.
- Non-reference processed pseudogenes represent a valuable resource for population genetics and forensic DNA testing.
More Related Videos
05:51A Strategy to Identify de Novo Mutations in Common Disorders such as Autism and Schizophrenia
Published on: June 15, 2011
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Exon Recombination
Exon shuffling follows “splice frame rules.” Each exon...
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Single Nucleotide Polymorphisms-SNPs
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Non-LTR Retrotransposons