SaVor - 一个可重复的结构变体呼叫和基准测试平台从短读数据
Trevor Mugoya1, Arun Sethuraman1
1Department of Biology, San Diego State University.
bioRxiv : the preprint server for biology
|December 11, 2025
概括
SaVor是一个新的工作流程,用于识别基因组中的结构变异 (SV),使用短读测序数据. 它提供了灵活性和可重复性,有助于分析基因组差异.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 结构变异 (SVs) 是基因组差异>1千基基对 (Kbp),源于DNA修复,基因组复制或可转移元素.
- 下一代测序 (NGS) 增加了短读基因组数据,需要高效的SV检测方法.
研究的目的:
- 引入SaVor,用于结构变化 (SV) 呼叫的灵活和可重复的工作流.
- 通过使用用户定义的合并参数,从短读序列数据生成共识的 SV 调用集.
主要方法:
- SaVor可以接受单或多车道短读配对终端Illumina数据或BAM文件.
- 该工作流被测试在1165个*Arabidopsis thaliana*全基因组序列上.
- 绩效与使用Lumpy.确定的SV相比进行了基准测试.
主要成果:
- 交叉呼叫 (由3个呼叫者支持的SV) 实现了高精度 (>0.91),但回忆率较低 (<0.51).
- 欧盟电话 (由至少1个呼叫者支持的SV) 显示了高回忆率 (>0.88),但精度较低 (<0.57).
- 合并策略影响了SV调用集的精度和召回之间的权衡.
结论:
- SaVor提供了一种可重现的方法,用于从短读数据调用SV.
- 下游分析需要仔细考虑基于所选择的合并策略的精密召回权衡.
- SaVor是一个开源的Snakemake管道,可在GitHub上使用.
更多相关视频
04:58Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance
Published on: December 13, 2024
3.9K
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
34.5K
相关概念视频
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Sanger Sequencing
772.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.8K
Next-generation Sequencing
97.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.6K
RACE - Rapid Amplification of cDNA Ends
7.1K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
7.1K
