Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Principles of Pharmacogenetics: Types of Genetic Variants01:27

Principles of Pharmacogenetics: Types of Genetic Variants

134
The human genome is over 99.9% identical between individuals, yet genetic differences exist at millions of bases. The human genome contains approximately 3 million variant positions per individual, many of which are heterozygous, contributing to genetic diversity and individual traits. Genetic variations include single-nucleotide polymorphisms (SNPs), insertions, deletions, and copy number variations (CNVs).SNPs, the most common variation, involve single-base changes in DNA. These can be...
134
Improving Translational Accuracy02:07

Improving Translational Accuracy

11.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.5K
Improving Translational Accuracy02:07

Improving Translational Accuracy

2.6K
2.6K
Genomics02:02

Genomics

35.1K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
35.1K
Pharmacogenomics: Identification of New Drug Targets01:29

Pharmacogenomics: Identification of New Drug Targets

113
Advances in genomics have profoundly influenced drug discovery by increasing both the speed and accuracy of pharmaceutical development. Pharmacogenomics, which examines how genetic variation influences drug response, facilitates the identification of novel therapeutic targets and enables patient stratification for personalized treatment. These strategies contribute to improved drug efficacy, minimized adverse effects, and more efficient clinical trial design.Mapping genetic differences...
113
Genome Annotation and Assembly03:36

Genome Annotation and Assembly

16.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
16.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Understanding the determinants of public trust in the health care system in China: an analysis of a cross-sectional survey.

Journal of health services research & policy·2018
Same author

Adverse Childhood Experiences, Epigenetic Measures, and Obesity in Youth.

The Journal of pediatrics·2018
Same author

LncRNA UCA1 sponges miR-204-5p to promote migration, invasion and epithelial-mesenchymal transition of glioma cells via upregulation of ZEB1.

Pathology, research and practice·2018
Same author

International variations in trust in health care systems.

The International journal of health planning and management·2018
Same author

Toll-like receptor 9 negatively regulates pancreatic islet beta cell growth and function in a mouse model of type 1 diabetes.

Diabetologia·2018
Same author

Methylation in OTX2 and related genes, maltreatment, and depression in children.

Neuropsychopharmacology : official publication of the American College of Neuropsychopharmacology·2018

Related Experiment Video

Updated: Apr 24, 2026

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
11:35

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA

Published on: August 21, 2016

12.5K

Pre-training genomic language model with variants for better modeling functional genomics.

Tianyu Liu1,2,3, Xiangyu Zhang2, Jiecong Lin4

  • 1Interdepartmental Program in Computational Biology & Bioinformatics, Yale University, New Haven, CT USA.

NPJ Artificial Intelligence
|April 23, 2026
PubMed
Summary

UKBioBERT, a new DNA language model, improves gene expression prediction by learning from genetic variants. Combining it with sequence-to-function models enhances understanding of gene regulation and its impact on complex traits.

Keywords:
Computational biology and bioinformaticsGenetics

More Related Videos

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
09:34

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease

Published on: April 4, 2018

36.1K
A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

10.5K

Related Experiment Videos

Last Updated: Apr 24, 2026

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
11:35

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA

Published on: August 21, 2016

12.5K
Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
09:34

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease

Published on: April 4, 2018

36.1K
A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

10.5K

Area of Science:

  • Genomics
  • Computational Biology
  • Functional Genomics

Background:

  • Genomic language models (GLMs) learn from DNA sequences to create representations.
  • Sequence-to-function (S2F) models predict gene expression from sequence data.
  • Bridging GLMs and S2F models is crucial for individualized gene expression prediction.

Purpose of the Study:

  • Introduce UKBioBERT, a GLM pre-trained on UK BioBank genetic variants.
  • Enhance gene expression prediction by integrating UKBioBERT with S2F models.
  • Investigate the relationship between genetic variants and gene expression variations.

Main Methods:

  • Pre-trained UKBioBERT on large-scale DNA sequences from UK BioBank.
  • Developed UKBioFormer and UKBioZoi by combining UKBioBERT with Enformer and Borzoi.
  • Evaluated model performance on gene expression prediction and eQTL identification.

Main Results:

  • UKBioBERT embeddings identify gene functions and improve cell line gene expression prediction.
  • UKBioFormer and UKBioZoi show enhanced performance in predicting predictable gene expression levels.
  • UKBioFormer effectively links genetic variants to expression variations, aiding in silico mutation analysis and eQTL identification.

Conclusions:

  • Integrating GLMs and S2F models advances functional genomics research.
  • UKBioBERT provides valuable embeddings for understanding gene expression predictability.
  • The developed models offer improved generalization across cohorts and facilitate genetic variant-trait association studies.