SIMLINK Enables Accurate Variant Pathogenicity Prediction through Modeling the Gene-Variant-Feature Association

Hong-Dong Li1,2, Chenlu Wang1, Dongfang Yan1

  • 1School of Computer Science and Engineering, Central South University, Changsha, Hunan, 410083, P.R. China.

Summary

SIMLINK improves variant pathogenicity prediction by modeling gene-variant-feature associations using a knowledge graph and separating linear from nonlinear components. This approach outperforms existing methods for predicting missense and synonymous variants.

Related Concept Videos

Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
Pharmacogenomics: Identification of New Drug Targets01:29

Pharmacogenomics: Identification of New Drug Targets

Advances in genomics have profoundly influenced drug discovery by increasing both the speed and accuracy of pharmaceutical development. Pharmacogenomics, which examines how genetic variation influences drug response, facilitates the identification of novel therapeutic targets and enables patient stratification for personalized treatment. These strategies contribute to improved drug efficacy, minimized adverse effects, and more efficient clinical trial design.Mapping genetic differences...
Genetic Variation01:25

Genetic Variation

Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles, which...
Point and Frameshift Mutations01:30

Point and Frameshift Mutations

Point mutations are genetic alterations involving the change of a single nucleotide base pair in DNA. Depending on how the alteration affects protein synthesis, they can lead to various consequences.Point mutations fall into the following types:Silent mutations occur when a nucleotide change does not alter the amino acid sequence due to the redundancy of the genetic code. For instance, changing ACC to ACA still encodes threonine, leaving the protein function unaffected. This occurs because...