Related Experiment Video
Updated: Aug 31, 2025

05:22
Analyzing Multifactorial RNA-Seq Experiments with DiCoExpress
Published on: July 29, 2022
3.6K
Doppelgänger spotting in biomedical gene expression data
Li Rong Wang1, Xin Yun Choy1, Wilson Wen Bin Goh2,3,4
1School of Computer Science and Engineering, Nanyang Technological University, 60 Nanyang Drive, 637551, Singapore.
Iscience
|August 22, 2022
Summary
Doppelgänger effects inflate machine learning model performance. We introduce doppelgangerIdentifier software to identify these deceptive similarities in biomedical data, ensuring reliable model evaluation.
Area of Science:
- Bioinformatics
- Machine Learning
- Genomics
Background:
- Doppelgänger effects (DEs) in machine learning (ML) involve chance sample similarities inflating model performance.
- This leads to overconfidence in ML model deployability, with no current tools for DE identification.
- DEs pose a significant challenge in biomedical data analysis.
Purpose of the Study:
- To introduce doppelgangerIdentifier, a novel software suite for identifying and verifying DEs.
- To demonstrate the prevalence of DEs in diverse biomedical gene expression datasets.
- To provide guidelines for managing DEs and improving ML model evaluation.
Main Methods:
- Development and application of the doppelgangerIdentifier software suite.
- Analysis of multiple disease and data types to assess DE prevalence.
- Exploration of batch effect influences on DE identification sensitivity.
Main Results:
- Doppelgänger effects are pervasive across various biomedical gene expression datasets.
- The doppelgangerIdentifier software effectively identifies and verifies DEs.
- Batch effects can impact the sensitivity of DE identification algorithms.
Conclusions:
- Doppelgänger effects are a widespread issue in biomedical ML, requiring specific identification tools.
- DoppelgangerIdentifier provides a solution for detecting and managing DEs.
- Doppelgänger verification is crucial for establishing reliable ML model evaluation baselines and ensuring meaningful insights from data.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.9K
Gene Duplication and Divergence
6.3K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.3K

