Related Experiment Video
Updated: Mar 18, 2026

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes
Published on: November 7, 2025
The Doppelgänger Effect: Hidden Duplicates in Databases of Transcriptome Profiles
Levi Waldron1, Markus Riester1, Marcel Ramos1
1Affiliations of authors: City University of New York School of Public Health, New York, NY (LW, MRa); Novartis Institutes for BioMedical Research, Cambridge, MA (MRi); Department of Biostatistics and Computational Biology, Dana-Farber Cancer Institute/Harvard Medical School, Boston, MA (GP); Center for Cancer Research, Massachusetts General Hospital, Boston, MA (MB).
Duplicate cancer transcriptome data in public databases can skew research findings. We developed a method to detect these "doppelgänger" profiles, even across different technologies, to ensure accurate genomic data analysis.
Area of Science:
- Genomics
- Bioinformatics
- Cancer Research
Background:
- Whole-genome analysis of cancer specimens is increasingly common.
- Re-using and sharing specimens can lead to duplicate expression profiles in public databases.
- Undetected duplicates, termed "doppelgängers," can negatively impact data re-analysis and research reproducibility.
Purpose of the Study:
- To propose and validate a method for accurately matching duplicate cancer transcriptomes.
- To address the challenge of identifying duplicates when nucleotide-level sequence data are unavailable.
- To establish a standard procedure for screening transcriptome databases for potential duplication.
Main Methods:
- Developed a method to match duplicate cancer transcriptomes using expression profiles.
- Applied the method to diverse datasets including ovarian, breast, bladder, and colorectal cancer microarrays.
- Validated the method by matching microarray and RNA sequencing profiles from The Cancer Genome Atlas (TCGA).
Main Results:
- Identified probable duplicate profiles in over 50% of the analyzed studies.
- Duplicates were found across different continents, technologies, publication years, and within TCGA.
- The method effectively matched samples profiled by different microarray technologies and by microarray versus RNA sequencing.
Conclusions:
- A robust method for detecting duplicate cancer transcriptomes is essential for accurate genomic data analysis.
- Doppelgänger-checking should be integrated into standard procedures for combining multiple genomic datasets.
- The doppelgangR Bioconductor package is provided to facilitate the screening of transcriptome databases for duplicates.
More Related Videos
Related Concept Videos
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Duplication of Chromatin Structure
The basic unit of the chromatin is the nucleosome, consisting of DNA wrapped around octameric histone proteins and short stretches of linker DNA separating individual nucleosomes. The histone proteins within the nucleosome have their...
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Ribosome Profiling
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...

