VarSCAT:一个计算工具,用于对基因组变异的序列上下文注释
Ning Wang1,2, Sofia Khan1, Laura L Elo1,2,3
1Turku Bioscience Centre, University of Turku and Åbo Akademi University, Turku, Finland.
PLoS computational biology
|August 11, 2023
概括
变异序列上下文注释工具 (VarSCAT) 解决了分析基因组变异序列上下文的局限性. VarSCAT提供多功能注释,用于并列重复和插入/删除,改善变体解释.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
背景情况:
- 基因组变异序列背景对于生物解释和变异调用至关重要.
- 目前用于注释各种序列上下文的现有方法,如并列重复和模糊的断点,是有限的.
研究的目的:
- 引入变异序列上下文注释工具 (VarSCAT),用于全面的基因组变异注释.
- 为分析变量序列上下文提供一个多功能和可定制的解决方案,包括并列重复和indels.
主要方法:
- 开发VarSCAT,这是一个用于注释基因组变异序列背景的工具.
- 人类变异组的注释,以评估并联重复区域中的断点模两可和变异特征.
主要成果:
- 与STR和indel注释的现有方法相比,VarSCAT显示出更大的多功能性和可定制性.
- 超过75%的人类生殖系和临床相关的基因表现出断点模两可.
- 在STR区域中,超过80%的人类生殖系小变异是indels,其大小与STR图案大小相关.
结论:
- VarSCAT增强了基因组变异的分析,特别是在STR等复杂区域和具有模两可的断点的区域.
- 这些发现突出了STR区域内人类生殖系变异中断点模两可的普遍性和indel特征.
更多相关视频
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.8K
09:10A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
Published on: May 22, 2018
9.2K
相关概念视频
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Genomics
36.5K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.5K
Single Nucleotide Polymorphisms-SNPs
15.3K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.3K
Cis-regulatory Sequences
3.0K
3.0K
Multi-species Conserved Sequences
4.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.0K
Genome-wide Association Studies-GWAS
13.6K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.6K
