CNV-Finder:简化复制数变异发现方法
Nicole Kuznetsov1,2, Kensuke Daida1, Mary B Makarious1,2,3
1Center for Alzheimer's and Related Dementias (CARD), National Institute on Aging and National Institute of Neurological Disorders and Stroke, National Institutes of Health, Bethesda, MD, USA.
bioRxiv : the preprint server for biology
|November 28, 2024
概括
新型深度学习管道CNV-Finder使用数组数据准确识别副本数变异 (CNV),用于神经疾病研究. 这种可扩展的工具加速了大规模的CNV检测,减少了手工工作量,并使进一步分析的样本优先级能够高效地确定.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 副本数变异 (CNVs) 在复杂的疾病病因和人口变异性方面至关重要.
- 精确的CNV检测对于疾病遗传学研究至关重要,通常需要大样本大小.
- 通过Log R比 (LRR) 和B基因基因频率 (BAF),Illumina基因型阵列提供了具有成本效益的CNV检测.
研究的目的:
- 开发和验证CNV-Finder,这是一个新的深度学习管道,用于大规模的CNV识别.
- 为了加快对样品的优先排序,用于随后的全基因组测序分析.
- 提高与神经疾病相关的基因中CNV检测的准确性和效率.
主要方法:
- 深度学习的整合,特别是长短期记忆 (LSTM) 网络,与数组数据 (LRR和BAF) 的整合.
- 在专家注释的样本和不同队列的验证上培训模型 (例如,全球帕金森遗传学计划).
- 开发一个交互式 Web 应用程序,用于可视化,审查和过 CNV 预测.
主要成果:
- CNV-Finder准确地检测删除和重复,在不同的队列中证明了有效性.
- 管道成功地识别了不同稀疏度,噪声和大小的区域的CNV,包括像17q21.31.1.这样的复杂区域.
- 人类反的整合提高了模型的性能,并降低了假阳性率.
结论:
- CNV-Finder是一个可扩展的,公开可用的资源,用于有效和准确的CNV识别.
- 该管道显著减少了研究人员的手工工作量,促进了有针对性的验证和下游分析.
- 背景理解和人类专业知识是提高CNV识别精度的关键.
相关概念视频
Comparing Copy Number Variations and SNPs
17.3K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.3K
Single Nucleotide Polymorphisms-SNPs
14.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.0K


