Related Experiment Video
Updated: Sep 12, 2025

In Vivo Functional Study of Disease-associated Rare Human Variants Using Drosophila
Published on: August 20, 2019
varCADD: large sets of standing genetic variation enable genome-wide pathogenicity prediction
Lusiné Nazaretyan1, Philipp Rentzsch1,2, Martin Kircher3,4
1Berlin Institute of Health at Charité - Universitätsmedizin Berlin, Berlin, 10117, Germany.
New datasets from human standing variation improve machine learning models for genetic variant prioritization. These larger, less biased datasets enhance accuracy, especially in challenging genomic regions, advancing genetic research.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Machine learning (ML) and artificial intelligence (AI) are crucial for identifying disease-causing genetic variations.
- Current ML models are limited by small, biased training datasets, often focusing only on protein-coding genes.
- Unbiased, comprehensive datasets are needed for accurate variant effect prediction.
Purpose of the Study:
- To develop improved training datasets for ML-based variant prioritization.
- To leverage human standing variation for creating larger and less biased training sets.
- To enhance the accuracy of predicting genetic variant deleteriousness.
Main Methods:
- Utilized whole-genome sequences from 71,156 individuals (gnomAD v3.0) to create training sets.
- Defined benign variants using frequent standing variation and deleterious variants using rare/singleton variation.
- Trained new models using the Combined Annotation Dependent Depletion (CADD) framework (v1.6).
Main Results:
- Alternative models achieved state-of-the-art accuracy, comparable to CADD v1.6/v1.7.
- Models trained on standing variation outperformed existing methods in specific genomic regions.
- Larger datasets captured broader genomic regions and rare annotations, including regulatory elements.
Conclusions:
- Human standing variation provides a powerful resource for training accurate genome-wide variant prioritization models.
- The proposed datasets are larger, less biased, and cover more genomic diversity than conventional sets.
- These datasets and models are publicly available to advance genetic research and applications.
More Related Videos
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Single Nucleotide Polymorphisms-SNPs
Evolutionary Relationships through Genome Comparisons
Genetic Variation
Genes exist in different versions called alleles,...
Modern Molecular Taxonomy
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...