ChimeraMiner: An Improved Chimeric Read Detection Pipeline and Its Application in Single Cell Sequencing

Na Lu1, Junji Li2, Changwei Bi3

  • 1State Key Lab of Bioelectronics, School of Biological Science and Medical Engineering, Southeast University, Nanjing 210096, China. nlu@seu.edu.cn.

Insights

ChimeraMiner is a new pipeline that efficiently detects chimeric reads from multiple displacement amplification (MDA) single-cell sequencing data. It significantly improves structural variation detection by removing false positives caused by these artifacts.

Area of Science:

  • Genomics
  • Bioinformatics
  • Molecular Biology

Background:

  • Multiple displacement amplification (MDA) is a widely used whole genome amplification technique for single-cell studies.
  • MDA, while effective, generates chimeric reads that disrupt downstream analyses.
  • Accurate detection of these chimeric sequences is crucial for reliable single-cell genomics.

Purpose of the Study:

  • To develop and evaluate ChimeraMiner, an improved pipeline for detecting and classifying chimeric reads from MDA sequencing data.
  • To compare the efficiency and performance of ChimeraMiner against existing methods.
  • To assess the impact of chimera removal on structural variation detection in single-cell datasets.

Main Methods:

  • Construction of the ChimeraMiner pipeline for analyzing MDA sequencing data.
  • Classification of chimeric sequences using ChimeraMiner.
  • Evaluation using two MDA datasets (MDA1 and MDA2) and comparison with a previous pipeline.
  • Application to single-cell datasets to assess structural variation detection.

Main Results:

  • ChimeraMiner demonstrated significantly improved processing efficiency, using only 43.4% of the time compared to the previous pipeline.
  • The pipeline accurately classified millions of chimeric read pairs, identifying a substantial number of novel chimeras missed by prior methods.
  • Application of ChimeraMiner effectively removed 83.8% of false positive structural variations in single-cell datasets.

Conclusions:

  • ChimeraMiner offers enhanced efficiency and accuracy in detecting and classifying chimeric reads from MDA-based single-cell sequencing.
  • The pipeline effectively mitigates false positives in structural variation detection, improving data reliability.
  • ChimeraMiner is a promising tool for widespread adoption in single-cell sequencing research.

Related Concept Videos

Uncertainty in Measurement: Reading Instruments02:46

Uncertainty in Measurement: Reading Instruments

Counting is the type of measurement that is free from uncertainty, provided the number of objects being counted does not change during the process. Such measurements result in exact numbers. By counting the eggs in a carton, for instance, one can determine exactly how many eggs are there in the carton. Similarly, the numbers of defined quantities are also exact. For example, 1 foot is exactly 12 inches, 1 inch is exactly 2.54 centimeters, and 1 gram is exactly 0.001 kilograms. Quantities...
51.0K
Cis-regulatory Sequences02:02

Cis-regulatory Sequences

Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.6K
Cis-regulatory Sequences02:02

Cis-regulatory Sequences

4.1K
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.6K
Sanger Sequencing01:57

Sanger Sequencing

DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
773.9K