Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Proteins: From Genes to Degradation02:11

Proteins: From Genes to Degradation

12.0K
Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick.  Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA...
12.0K
Improving Translational Accuracy02:07

Improving Translational Accuracy

9.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.1K
Genome Annotation and Assembly03:36

Genome Annotation and Assembly

18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Genome Size and the Evolution of New Genes03:21

Genome Size and the Evolution of New Genes

2.4K
2.4K
Calmodulin-dependent Signaling01:16

Calmodulin-dependent Signaling

5.1K
Calmodulin (CaM) is a calcium-binding protein in eukaryotes that controls various calcium-regulated cellular processes. It has four calcium-binding sites that bind calcium to form the calcium-calmodulin ( Ca2+-CaM) complex. GPCR stimulation increases the calcium levels in the cells that bind to CaM and induces a conformational change.
The Ca2+-CaM complex does not have enzymatic activity by itself. Instead, the complex binds downstream target proteins, including membrane proteins or enzymes,...
5.1K
Leaky Scanning02:28

Leaky Scanning

5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA.  Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Pulmonary fibrosis after COVID-19 is characterized by airway abnormalities and elevated club cell secretory protein-16.

JCI insight·2026
Same author

Signal in the Noise: Polygenic Scores and the Problem of Defining Idiopathic Pulmonary Fibrosis.

American journal of respiratory and critical care medicine·2026
Same author

Progressive Fusion of Multi-Scale Mamba Context and Local Detail Priors for Infrared Small Target Detection.

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026
Same author

Delta-Front Cities as Dynamic Receptors of Microplastic Fluxes: Seasonal Source Switching across Multioutlet Deltas.

Environmental science & technology·2026
Same author

CpG Atlas: A centralized multi-layer database and AI interface for DNA methylation research.

bioRxiv : the preprint server for biology·2026
Same author

Performance of Age-Adjusted Whole Genome Sequencing Telomere Length in Idiopathic Pulmonary Fibrosis.

American journal of respiratory and critical care medicine·2026

Related Experiment Video

Updated: Jun 7, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

504

Cell2Sentence: Teaching Large Language Models the Language of Biology.

Daniel Levine1, Syed Asad Rizvi1, Sacha Lévy1

  • 1Department of Computer Science, Yale University, New Haven, CT, USA.

Biorxiv : the Preprint Server for Biology
|November 18, 2024
PubMed
Summary

Cell2Sentence (C2S) transforms gene expression data into "cell sentences," enabling large language models to understand single-cell biology. This method facilitates cell generation and accurate cell type annotation for diverse biological applications.

More Related Videos

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
09:34

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data

Published on: September 25, 2021

3.9K
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
10:41

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms

Published on: May 9, 2017

9.2K

Related Experiment Videos

Last Updated: Jun 7, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

504
A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
09:34

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data

Published on: September 25, 2021

3.9K
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
10:41

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms

Published on: May 9, 2017

9.2K

Area of Science:

  • Computational Biology
  • Bioinformatics
  • Genomics

Background:

  • Single-cell transcriptomics generates high-dimensional gene expression data.
  • Integrating advanced computational methods like natural language processing (NLP) can unlock new insights.
  • Existing NLP models require adaptation for biological data.

Purpose of the Study:

  • To introduce Cell2Sentence (C2S), a novel method for adapting large language models (LLMs) to single-cell transcriptomics.
  • To demonstrate the utility of C2S in enabling LLMs to perform biological tasks.
  • To bridge the gap between NLP and single-cell biology.

Main Methods:

  • Transforming gene expression data into "cell sentences."
  • Fine-tuning pre-trained LLMs (e.g., GPT-2) using cell sentences.
  • Evaluating the fine-tuned models on tasks like cell generation and cell type annotation.

Main Results:

  • Fine-tuned GPT-2 models can generate biologically valid cells based on cell type inputs.
  • The models accurately predict cell types from cell sentences.
  • LLMs fine-tuned with C2S acquire a significant understanding of single-cell biology.

Conclusions:

  • Cell2Sentence (C2S) provides a flexible framework for integrating NLP with transcriptomics.
  • C2S enables LLMs to perform complex biological tasks, including cell generation and annotation.
  • This approach leverages existing NLP models and libraries for broad biological applications.