相关实验视频
Updated: Jun 12, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.0K
探索序列背景对SNP基因型调用错误的影响,使用基于AI的自编码器方法调用全基因组测序数据
Krzysztof Kotlarz1, Magda Mielczarek1, Przemysław Biecek2,3
1Biostatistics Group, Department of Genetics, Wroclaw University of Environmental and Life Sciences, Wroclaw 51-631, Poland.
NAR genomics and bioinformatics
|September 25, 2024
概括
这项研究确定了整个基因组测序数据中不正确的单核酸多态化 (SNP) 调用中的系统模式. 了解这些变异调用错误可以提高基因组数据的准确性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 动物遗传学动物遗传学
背景情况:
- 变异调用对于全基因组测序 (WGS) 数据分析至关重要,但容易出现错误.
- 错误的单核酸多态化 (SNP) 调用可能会影响下游基因组分析.
- 识别变量调用错误中的模式对于提高数据质量至关重要.
研究的目的:
- 调查SNP错误调用和变量质量指标之间的关联.
- 使用核酸背景识别SNP调用错误中的系统模式.
- 开发用于检测和潜在地纠正WGS数据中的错误SNP调用方法.
主要方法:
- 从WGS (Illumina NovaSeq 6000) 的SNP基因型与Holstein-Friesian牛的基因型微阵列 (EuroGMD50K) 的基因型基因型相比较.
- 定义了正确的 (666,333个SNP) 和不正确的 (4,557个SNP) SNP集.
- 采用了自动编码器,一类支持向量机器和隔离森林算法来识别系统错误.
主要成果:
- 大约59.53%的错误SNP表现出系统的模式,而剩下的则是随机错误.
- 序列上下文"CGC"经常与错误地称为"C"变体有关.
- 错误的"T"而不是"A"调用与下游的"T"核酸有关.
结论:
- 在SNP调用中存在系统性错误,并且与特定的序列环境和核酸标记模式有关.
- 这些发现为WGS数据中变量调用错误的来源提供了洞察力.
- 更好地了解错误模式可以导致牛和其他物种更准确的基因组分析.
相关概念视频
Genome Copying Errors
4.2K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.2K
Single Nucleotide Polymorphisms-SNPs
14.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.8K
Comparing Copy Number Variations and SNPs
17.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.6K
Sanger Sequencing
753.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
753.8K
Genome-wide Association Studies-GWAS
13.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.2K
Next-generation Sequencing
88.4K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.4K

