Related Experiment Video
Updated: May 31, 2026

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
Published on: March 22, 2018
DupyliCate: mining, classifying, and characterizing gene duplications
Shakunthala Natarajan1, Boas Pucker2
1Plant Biotechnology and Bioinformatics, Institute for Cellular and Molecular Botany, University of Bonn, Kirschallee 1, 53115, Bonn, Germany.
Abstract:
Paralogs, copies of a gene, form an important basis for novelty during evolution. Analysis of such gene duplications is important to understand the emergence of novel traits during evolution. DupyliCate is a Python tool that has been developed for this purpose. With the ability to process multiple datasets concurrently, flexible features, and parameters to set species-specific thresholds, DupyliCate offers a high-throughput method for gene copy identification and analysis. The different available parameters and modes are explored in detail based on Arabidopsis thaliana datasets. Proof of concept for the tool is presented by characterizing well known duplications in different plants, and its broad applicability is demonstrated by running it on diverse datasets including complex plant genome sequences with high heterozygosity. Further, two case studies involving the evolution of FLAVONOL SYNTHASE (FLS) genes in Brassicales, and the evolution of flavonol synthesis regulating myeloblastosis (MYB) transcription factors-MYB12 and MYB111 across a large number of plant species, are presented as exemplar use cases. The tool's applicability beyond plants is demonstrated on Escherichia coli, Saccharomyces cerevisiae, and Caenorhabditis elegans datasets. DupyliCate is available at: https://github.com/ShakNat/DupyliCate .
Related Concept Videos
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are characterized.
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Evolutionary Relationships through Genome Comparisons
Exon Recombination
Exon shuffling follows “splice frame rules.” Each exon has three reading...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Chromosome Duplication
The basic unit of the chromatin is the nucleosome, consisting of DNA wrapped around octameric histone proteins and short stretches of linker DNA separating individual nucleosomes. The histone proteins within the nucleosome have their...

