Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics
Gonzalo Benegas1, Gökcen Eraslan2, Yun S Song1,3,4
1Computer Science Division, University of California, Berkeley.
Biorxiv : the Preprint Server for Biology
|February 24, 2025
Summary
TraitGym, a new dataset, aids in identifying causal genetic variants for diseases. It benchmarks machine learning models, revealing strengths of alignment-based and functional genomics approaches for Mendelian and complex traits.
Area of Science:
- Genomics
- Machine Learning
- Computational Biology
Background:
- Machine learning (ML) is crucial for identifying causal genetic variants in Mendelian and complex traits.
- Current ML approaches include supervised sequence-to-function and self-supervised DNA language models.
- A lack of curated datasets with accurate labels hinders benchmarking, especially for non-coding variants.
Purpose of the Study:
- Introduce TraitGym, a curated dataset of regulatory genetic variants for benchmarking ML models.
- Evaluate the performance of various ML models in predicting causal variants.
- Provide insights into the capabilities and limitations of different prediction strategies.
Main Methods:
- Developed TraitGym, a dataset of causal/candidate regulatory variants and controls for 113 Mendelian and 83 complex traits.
- Framed variant prediction as a binary classification task.
- Benchmarked supervised, self-supervised, hybrid, and ensemble ML models.
Main Results:
- Alignment-based models (CADD, GPN-MSA) performed well for Mendelian and complex disease traits.
- Functional genomics models (Enformer, Borzoi) excelled in complex non-disease traits.
- Evo2 showed scalability benefits but lagged behind alignment models, especially for enhancer variants.
Conclusions:
- TraitGym facilitates comprehensive benchmarking of ML models for causal variant identification.
- Different ML approaches show varying performance depending on trait type and variant class.
- The dataset and benchmark provide valuable resources for advancing genetic variant interpretation.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.6K
Cis-regulatory Sequences
9.8K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
9.8K
Genome-wide Association Studies-GWAS
13.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.2K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Gene Evolution - Fast or Slow?
7.0K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.0K
Incomplete Dominance
22.0K
Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
22.0K


