GraphSlimmer:保留读取可映射性,使用最小数量的变体
Neda Tavakoli1, Daniel Gibney2, Srinivas Aluru1
1School of Computational Science and Engineering, Georgia Institute of Technology, Atlanta, Gxeorgia, USA.
概括
这项研究提出了一种方法,可以显著减少基因组变异,同时确保准确的读取映射. 该技术保留了必不可少的哈普洛型信息,使得大型基因组数据集的有效分析成为可能.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 现代基因组数据集 (例如1000基因组项目) 包含数以百万计的变异,这给读取映射等任务带来了计算挑战.
- 完整的变异集可能会降低映射精度,并且由于庞大的基因组图数据结构,需要大量的计算资源.
研究的目的:
- 开发一种用于识别基因组变异的最小子集的技术.
- 为了确保所有长度高达指定长度的子字符串 (α) 仍然可以与使用缩小变量设置的有限的哈明或编辑距离 (δ) 保持对齐.
主要方法:
- 证明了变量子集选择优化问题的NP硬度和不可接近性.
- 开发了整数线性编程 (ILP) 公式用于变量选择,包括一个编辑距离公式,该公式将问题分解为变量位置.
- 缩放编辑距离ILP配方,以处理来自1000基因组项目的22号染色体的所有变异.
主要成果:
- 表明可以实现显著减少变体.
- 对于中等长度的阅读 (α=1000),可以删除超过75%的变体,同时保持阅读可映射性,编辑距离最多为1.5.
- 拟议的ILP配方成功扩展到大规模的基因组数据.
结论:
- 一种最小变异子集选择技术可以大大降低基因组分析的计算负担.
- 该方法保留了精确读取映射的基本信息,解决了使用完整变体集的局限性.
- 这种方法为分析大规模基因组数据集提供了一个计算效率高的替代方案.
相关概念视频
Conserved Binding Sites
1.7K
1.7K
Improving Translational Accuracy
2.6K
2.6K
Single Nucleotide Polymorphisms-SNPs
15.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.0K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Mismatch Repair
4.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
4.8K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K


