FASTQ和对齐阅读顺序对结构变异调用从长时间阅读测序数据的影响
Kyle J Lesack1,2, James D Wasmuth1,2
1Faculty of Veterinary Medicine, University of Calgary, Calgary, Alberta, Canada.
PeerJ
|March 19, 2024
概括
从长读序列数据调用结构变体 (SV) 对FASTQ读序敏感,影响可重现性. 研究人员和工具开发人员必须考虑这种输入顺序灵敏度,以准确检测和一致的结果.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
背景情况:
- 来自DNA测序数据的结构变异 (SV) 调用面临着像对齐模两可和缺乏基准数据集这样的挑战.
- 虽然呼叫者选择和参数影响SV呼叫,但FASTQ读取顺序对长读数据的影响仍然未被探索.
研究的目的:
- 为了评估不同SV呼叫者对FASTQ在长读序列数据中读取顺序的灵敏度.
- 确定导致在SV调用中读取订单灵敏度的因素.
主要方法:
- 使用了PacBio DNA测序数据,这些数据来自*Caenorhabditis elegans*和*Arabidopsis thaliana*.
- 通过使用各种SV调用器和对齐器,从原始和换的FASTQ文件生成的SV调用进行了比较.
- 分析了测序深度和对齐分类算法的影响.
主要成果:
- 在不同的呼叫者和物种中,FASTQ读数顺序显著影响了SV预测.
- pbsv呼叫者对读取顺序表现出高灵敏度,在高序列深度下超过70%的异议.
- SAMtools对齐分类被确定为变化源.
结论:
- 从长时间读取数据呼叫的SV对FASTQ读取顺序敏感,这是以前未被识别的因素.
- 这种敏感性影响研究复制和开发一致的SV调用协议.
- 研究人员和工具开发人员应该解决输入顺序的敏感性,以改进 SV 检测和基准测试.
更多相关视频
04:58Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance
Published on: December 13, 2024
2.4K
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.7K
相关概念视频
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K
Next-generation Sequencing
88.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.8K
Sanger Sequencing
754.3K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.3K
Maxam-Gilbert Sequencing
11.2K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.2K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
