Related Experiment Video
Updated: Sep 14, 2025

06:56
Purification of High Molecular Weight Genomic DNA from Powdery Mildew for Long-Read Sequencing
Published on: March 31, 2017
11.8K
Targeted decontamination of sequencing data with CLEAN.
Marie Lataretu1,2, Sebastian Krautwurst2, Matthew R Huska1
1Genome Competence Center, Robert Koch Institute, 13353 Berlin, Germany.
NAR Genomics and Bioinformatics
|July 25, 2025
Summary
We developed CLEAN, a pipeline to remove unwanted sequences like spike-ins and host DNA from sequencing data. This tool ensures cleaner data for faster, more accurate genomic and transcriptomic analyses.
Area of Science:
- Genomics
- Transcriptomics
- Bioinformatics
Background:
- Biological and medical research relies heavily on sequence data analysis.
- Sequence data collections often contain contaminants such as artificial spike-ins, overrepresented ribosomal RNA (rRNA), and human host DNA.
- Existing methods for removing these contaminants are often not fully reproducible or traceable.
Purpose of the Study:
- To develop a robust and reproducible pipeline for removing unwanted sequences from sequencing data.
- To address the challenge of overlooked contaminants, including technology-specific spike-ins and human DNA.
- To improve the accuracy and efficiency of downstream genomic and transcriptomic analyses.
Main Methods:
- Developed the CLEAN pipeline for processing both long- and short-read sequencing data.
- The pipeline is designed to remove technology-specific control sequences (e.g., Illumina, Nanopore), human host DNA, and rRNA.
- CLEAN generates purified sequences and provides statistics on identified contaminants in a summary report.
Main Results:
- CLEAN effectively removes unwanted sequences, including spike-ins and host DNA, from various sequencing datasets.
- The pipeline supports platform-independent data analysis for both genomics and transcriptomics.
- Output data is ready for immediate use in subsequent analyses, leading to faster computations and improved results.
Conclusions:
- CLEAN offers a reproducible and traceable solution for sequence data decontamination.
- The pipeline enhances the reliability of biological and medical research by ensuring data integrity.
- CLEAN is freely available, promoting standardized and efficient data processing in the scientific community.
Related Concept Videos
DNA Isolation
40.5K
DNA isolation protocols can be fast and straightforward or complex and time-consuming depending on the type and quality of DNA required for further processing. For example, plasmid DNA extraction is a bit more complicated than genomic DNA extraction because of the need for an appropriate lysis method to separate plasmid DNA from gDNA during isolation. However, for specific applications, such as long-range DNA sequencing that require a good yield of high- quality DNA samples, we need to follow...
40.5K
Maxam-Gilbert Sequencing
11.5K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.5K

