Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genomics02:02

Genomics

35.8K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
35.8K
Genome Annotation and Assembly03:36

Genome Annotation and Assembly

18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Germline hypomethylation shapes dynamic CpG reservoirs in ape genomes.

bioRxiv : the preprint server for biology·2026
Same author

Biological functions of BAF57, its role in disease pathogenesis, and treatment: From molecular mechanisms to clinical translation.

Progress in biophysics and molecular biology·2026
Same author

Uncertainty-aware synthetic lethality prediction with pretrained foundation models.

bioRxiv : the preprint server for biology·2026
Same author

An integrated view of the structure and function of the human 4D nucleome.

Nature·2025
Same author

MIMYR: Generative modeling of missing tissue in spatial transcriptomics.

bioRxiv : the preprint server for biology·2025
Same author

TissueNarrator: Generative Modeling of Spatial Transcriptomics with Large Language Models.

bioRxiv : the preprint server for biology·2025

Related Experiment Video

Updated: Jun 4, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

500

L2G: Repurposing Language Models for Genomics Tasks.

Wenduo Cheng1, Junhong Shen2, Mikhail Khodak3

  • 1Ray and Stephanie Lane Computational Biology Department, School of Computer Science, Carnegie Mellon University, Pittsburgh, PA 15213, USA.

Biorxiv : the Preprint Server for Biology
|December 23, 2024
PubMed
Summary

Repurposing large language models (LLMs) for genomics bypasses data and compute challenges. The L2G method adapts LLMs for genomic tasks, achieving superior performance without extensive DNA pre-training.

More Related Videos

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

9.1K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

637

Related Experiment Videos

Last Updated: Jun 4, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

500
A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes
09:10

A Fast and Quantitative Method for Post-translational Modification and Variant Enabled Mapping of Peptides to Genomes

Published on: May 22, 2018

9.1K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

637

Area of Science:

  • Genomics
  • Bioinformatics
  • Computational Biology

Background:

  • Foundation models (FMs) are transforming genomics, mirroring successes in natural language processing (NLP).
  • Developing genomic FMs from scratch is computationally expensive and requires vast, high-quality datasets.
  • Large language models (LLMs) in NLP benefit from industrial-scale data and infrastructure.

Purpose of the Study:

  • To adapt existing LLMs for genomics, overcoming data and computational bottlenecks.
  • To introduce L2G, a method for repurposing LLMs for diverse genomic applications.
  • To evaluate the efficacy of LLM adaptation in genomics.

Main Methods:

  • Leveraging cross-modal transfer from NLP transformers to genomic data.
  • Employing neural architecture search (NAS) to adapt LLM architectures.
  • Utilizing a novel three-stage training procedure for genomic tasks.

Main Results:

  • L2G achieves superior performance on over half of tested genomics benchmark tasks.
  • The model outperforms fine-tuned genomic FMs and task-specific models.
  • L2G successfully identifies significant transcription factor motifs in enhancer activity prediction.

Conclusions:

  • Pre-trained language models demonstrate remarkable generalizability to out-of-domain tasks like genomics.
  • L2G offers an efficient, less resource-intensive approach to developing genomic models.
  • This work opens new avenues for leveraging LLMs in genomic research.