Related Experiment Video
Updated: Oct 7, 2025

Author Spotlight: Transmitochondrial Cybrid Generation Using Cancer Cell Lines
Published on: March 17, 2023
mitoDataclean: A machine learning approach for the accurate identification of cross-contamination-derived tumor
Liping Su1, Shanshan Guo1, Wenjie Guo1
1State Key Laboratory of Cancer Biology and Department of Physiology and Pathophysiology, Fourth Military Medical University, Xi'an, China.
Abstract:
Next-generation sequencing (NGS) of mitochondrial DNA (mtDNA) has widespread applications in aging and cancer studies. However, cross-contamination of mtDNA constitutes a major concern. Previous methods for the detection of mtDNA contamination mainly focus on haplogroup-level phylogeny, but neglect haplotype-level differences, leading to limited sensitivity and accuracy. In our study, we present mitoDataclean, a random-forest-based machine learning package for accurate identification of cross-contamination, evaluation of contamination levels and detection of contamination-derived variants in mtDNA NGS data. Comprehensive optimization of mitoDataclean revealed that training simulation with mixtures of small haplogroup distance and low polymorphic difference was critical for optimal modeling. Compared to existing methods, mitoDataclean exhibited significantly improved sensitivity and accuracy for the detection of sample contamination in simulated data. In addition, mitoDataclean achieved area under the curve values of 0.91 and 0.97 for discerning genuine and contamination-derived mtDNA variants in a simulated Western dataset and private sequencing contamination data, respectively, suggesting that this tool may be applicable for different populations and samples with different sources of contamination. Finally, mitoDataclean was further evaluated in several private and public datasets and showed a robust ability for contamination detection. Altogether, our study demonstrates that mitoDataclean may be used for accurate detection of contaminated samples and contamination-derived variants in mtDNA NGS data.
Insights
mitoDataclean accurately identifies mitochondrial DNA (mtDNA) contamination in next-generation sequencing (NGS) data. This machine learning tool improves sensitivity and accuracy for detecting contaminated samples and variants, crucial for aging and cancer research.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Next-generation sequencing (NGS) of mitochondrial DNA (mtDNA) is vital for aging and cancer research.
- mtDNA cross-contamination is a significant challenge, limiting accuracy in current studies.
- Existing detection methods lack sensitivity by focusing on haplogroup-level analysis, ignoring haplotype variations.
Purpose of the Study:
- To introduce mitoDataclean, a novel machine learning package for detecting mtDNA contamination in NGS data.
- To evaluate the accuracy and sensitivity of mitoDataclean compared to existing methods.
- To assess the capability of mitoDataclean in identifying contamination-derived variants.
Main Methods:
- Development of a random-forest-based machine learning package, mitoDataclean.
- Optimization of training simulations using mixtures with small haplogroup distance and low polymorphic difference.
- Validation using simulated datasets, private sequencing contamination data, and public datasets.
Main Results:
- mitoDataclean demonstrated significantly improved sensitivity and accuracy in detecting simulated mtDNA contamination.
- Achieved high area under the curve values (0.91 and 0.97) for distinguishing genuine from contamination-derived mtDNA variants.
- Showcased robust contamination detection capabilities across diverse private and public datasets.
Conclusions:
- mitoDataclean offers a sensitive and accurate solution for identifying mtDNA contamination in NGS data.
- The tool effectively detects contamination-derived variants, enhancing data reliability.
- mitoDataclean is applicable across different populations and contamination sources, supporting aging and cancer research.

