nf-core/pacvar:用于分析长期读取的PacBio全基因组和重复扩张测序数据的管道
Tanya Jain1, Claire Clelland1,2
1Weill Institute for Neurosciences, University of California, San Francisco, CA, 94158, United States.
Bioinformatics (Oxford, England)
|March 18, 2025
概括
一个新的生物信息管道,nf-core/pacvar,处理PacBio长时间读取的测序数据,用于全基因组测序和有针对性的重复分析. 这使得能够有效地调用变体,并对神经退行性疾病相关的重复区域进行表征.
科学领域:
- 基因组学和生物信息学
- 分子生物学分子生物学
背景情况:
- 太平洋生物科学 (PacBio) 长读测序为全基因组注释和表征复杂的重复区域提供了优势,特别是与神经退行性疾病相关的复杂重复区域.
- 长读全基因组测序 (WGS) 有助于检测短读测序经常错过的结构变异.
- 原始测序数据 (二进制对齐图) 需要在全面分析之前进行广泛的处理,突出需要简化生物信息解决方案.
研究的目的:
- 开发和介绍nf-core/pacvar,一个直观和全面的生物信息管道.
- 为了能够分析PacBio单分子PureTarget和WGS数据.
- 为了促进变体调用和重复表征的测序数据的快速,端到端和并行处理.
主要方法:
- nf-core/pacvar管道的开发是为了去多重化和并行化预处理步骤.
- 它集成了PacBio长读数据的变量调用和重复表征.
- 管道的设计是为了最小的配置和很少的依赖.
主要成果:
- nf-core/pacvar为分析PacBio PureTarget和WGS数据提供了一个全面的解决方案.
- 管道能够实现快速,并行和端到端的处理.
- 它有效地处理预处理,变量调用和重复表征.
结论:
- nf-core/pacvar显著简化了PacBio长时间读取的测序数据的分析.
- 该管道支持重复区域和结构变异的高效表征.
- 它是基因组学研究的宝贵工具,特别是用于神经退行性疾病研究.
相关概念视频
Sanger Sequencing
752.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
752.0K
Next-generation Sequencing
86.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
86.7K
Comparing Copy Number Variations and SNPs
16.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
16.9K
RNA-seq
9.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.7K
RACE - Rapid Amplification of cDNA Ends
6.3K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
6.3K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K


