Fastq-dupaway:一个快速和内存高效的工具,用于单端和配对NGS数据的脱复制
A I Sigorskikh1, M A Kompaniets2, I S Ilnitskiy1,3
1Faculty of Bioengineering and Bioinformatics, Lomonosov Moscow State University, Moscow, 119234, Russia.
Scientific reports
|November 25, 2025
概括
Fastq-dupaway有效地从下一代测序 (NGS) 数据中删除聚合酶连锁反应 (PCR) 重复物. 该工具使用最小的RAM,使个人计算机上大规模的数据处理成为可能.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 下一代测序 (NGS) 产生了庞大的数据集,需要高效的预处理.
- 聚合酶连锁反应 (PCR) 重复删除对于减少NGS数据中的放大偏差至关重要.
- 现有的 de novo 脱复制工具需要大量的计算资源,特别是高RAM,阻碍了大规模的数据分析.
研究的目的:
- 介绍Fastq-dupaway,这是一个用于有效地从NGS数据中删除PCR重复的新工具.
- 解决大型数据集当前脱复制方法的计算局限性.
- 允许在标准个人电脑上处理大量的测序数据.
主要方法:
- Fastq-dupaway在主要模式中运行,该模式是为低RAM使用量 (2-10GB) 而设计的,独立于输入数据大小.
- 该工具在磁盘空间中需要大约是输入文件大小的2倍.
- 支持单端和双端测序数据.
主要成果:
- Fastq-dupaway实现了与现有的 de novo 脱复制工具相比,可比或高达三倍的快速处理速度.
- 在重复删除中保持高准确度.
- 在标准硬件上展示了大型NGS数据集 (100GB+) 的高效处理.
结论:
- 在大型NGS数据中,Fastq-dupaway为PCR重复删除提供了一个计算效率高的解决方案.
- 该工具通过减少硬件要求,使大规模测序数据集的分析民主化.
- 通过提供高质量,偏差减少的测序数据,促进下游分析.
相关概念视频
Sanger Sequencing
772.9K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.9K
Next-generation Sequencing
97.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.6K
RNA-seq
11.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.7K
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Gene Duplication and Divergence
7.8K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
7.8K


