varCADD:大量的常态遗传变异使全基因组病原性预测成为可能
Lusiné Nazaretyan1, Philipp Rentzsch1,2, Martin Kircher3,4
1Berlin Institute of Health at Charité - Universitätsmedizin Berlin, Berlin, 10117, Germany.
来自人类站立变异的新数据集改善了用于遗传变异优先级的机器学习模型. 这些更大,更不偏的数据集提高了准确性,特别是在具有挑战性的基因组区域,推进了遗传研究.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 机器学习 (ML) 和人工智能 (AI) 对于识别引起疾病的遗传变异至关重要.
- 目前的ML模型受限于小的,有偏见的训练数据集,通常只关注蛋白质编码基因.
- 需要公正,全面的数据集来准确预测变异效应.
研究的目的:
- 开发改进的培训数据集,以基于机器学习的变体优先级.
- 为了利用人类身高的变化来创建更大,更少偏见的训练集.
- 为了提高预测遗传变异有害性的准确性.
主要方法:
- 利用来自71,156个个体的全基因组序列 (gnomAD v3.0) 来创建训练集.
- 定义的良性变异使用频繁的站立变异和有害变异使用罕见/单个变异.
- 使用Combined Annotation Dependent Depletion (CADD) 框架 (v1.6) 训练新的模型.
主要成果:
- 替代模型实现了最先进的准确性,与CADD v1.6 / v1.7.7相当.
- 在静态变异上训练的模型在特定的基因组区域中表现优于现有方法.
- 较大的数据集捕获了更广泛的基因组区域和罕见的注释,包括监管元素.
结论:
- 人类站立变异为训练精确的全基因组变异优先级模型提供了强大的资源.
- 拟议的数据集较大,不那么有偏见,并且比传统数据集覆盖更多的基因组多样性.
- 这些数据集和模型是公开可用的,以推进遗传研究和应用.
更多相关视频
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
相关概念视频
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Single Nucleotide Polymorphisms-SNPs
Evolutionary Relationships through Genome Comparisons
Genetic Variation
Genes exist in different versions called alleles,...
Modern Molecular Taxonomy
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
