Related Experiment Video
Updated: Mar 9, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.6K
Duplicates, redundancies and inconsistencies in the primary nucleotide databases: a descriptive study
Qingyu Chen1, Justin Zobel1, Karin Verspoor1
1Department of Computing and Information Systems, The University of Melbourne, Parkville, VIC, 3010, Australia.
Database : the Journal of Biological Databases and Curation
|January 13, 2017
Summary
Biological databases like GenBank contain numerous duplicates. This study quantifies duplicate types and their impact on data analysis, revealing inconsistencies in nucleotide sequence data.
Area of Science:
- Bioinformatics
- Genomics
- Database Management
Background:
- The International Nucleotide Sequence Database Collaboration (INSDC) comprises major nucleotide sequence databases: GenBank, EMBL, and DDBJ.
- These databases accumulate records over decades, leading to inherent duplicates, redundancies, and inconsistencies.
- Current duplicate detection methods lack comprehensive assessment and consistent assumptions, hindering effective management.
Purpose of the Study:
- To rigorously assess the scale, characteristics, and impact of duplicates within INSDC databases.
- To provide a quantitative analysis of duplicate types at both sequence and annotation levels.
- To demonstrate the practical consequences of data duplication on downstream analyses.
Main Methods:
- Retrospective analysis of merged groups within the INSDC databases.
- Utilized a benchmark dataset of manually identified duplicates (67,888 merged groups, 111,823 duplicate pairs across 21 organisms).
- Categorized duplicates by type (sequence and annotation) and assessed their impact through a case study on GC content and melting temperature.
Main Results:
- Different organisms exhibit varying prevalence of distinct duplicate types.
- Duplicates were categorized at both sequence and annotation levels, with quantitative statistics provided.
- A case study demonstrated that duplicates can lead to inconsistent results in tasks like GC content and melting temperature calculations.
Conclusions:
- The presence of duplicates in biological databases introduces significant redundancy and can compromise the accuracy of analyses.
- Understanding the prevalence and types of duplicates is crucial for improving data quality and reliability.
- This work provides a foundation for developing more effective strategies to identify and mitigate data duplication in large-scale biological sequence repositories.
Related Concept Videos
Gene Duplication and Divergence
8.1K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
8.1K
Comparing Copy Number Variations and SNPs
19.0K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.0K
Evolutionary Relationships through Genome Comparisons
7.1K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.1K
Multi-species Conserved Sequences
4.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.9K
Mismatch Repair
44.3K
Overview
44.3K
Mismatch Repair
6.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
6.8K

