DupChecker: a bioconductor package for checking high-throughput genomic data redundancy in meta-analysis
Quanhu Sheng, Yu Shyr, Xi Chen1
1Center for Quantitative Sciences, Vanderbilt University School of Medicine, Nashville, TN 37232, USA. xi.steven.chen@gmail.com.
Duplicate samples in genomic meta-analysis can lead to false positives. The DupChecker package efficiently identifies these duplicates using MD5 fingerprints, ensuring reliable biological signal detection.
Area of Science:
- Genomics
- Bioinformatics
Background:
- Meta-analysis is crucial for high-throughput genomic data analysis, enhancing the power to detect biological signals.
- Sample duplication in public databases, particularly for gene expression data, is a common issue.
- Unidentified duplicates can cause false positives, skewed clustering, and model overfitting.
Purpose of the Study:
- To introduce DupChecker, a Bioconductor package designed for efficient identification of duplicated samples in genomic datasets.
- To provide a practical tool for researchers to ensure data integrity before meta-analysis.
Main Methods:
- Development of the DupChecker package utilizing MD5 fingerprinting of raw data.
- Implementation of a robust method for detecting sample duplication.
Main Results:
- DupChecker efficiently identifies duplicated samples through MD5 fingerprint generation.
- A demonstration using real data showcases the package's functionality and output.
Conclusions:
- Failure to address sample duplication can compromise the validity of meta-analysis results.
- It is recommended to apply DupChecker to all gene expression datasets prior to analysis to prevent data contamination.
More Related Videos
04:58Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance
Published on: December 13, 2024
09:06High-throughput Identification of Gene Regulatory Sequences Using Next-generation Sequencing of Circular Chromosome Conformation Capture 4C-seq
Published on: October 5, 2018
