Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genome Annotation and Assembly03:36

Genome Annotation and Assembly

20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Cell Specific Gene Expression01:58

Cell Specific Gene Expression

5.4K
5.4K
Cell Specific Gene Expression01:58

Cell Specific Gene Expression

16.2K
Multicellular organisms contain a variety of structurally and functionally distinct cell types, but the DNA in all the cells originated from the same parent cells. The differences in the cells can be attributed to the differential gene expression. Liver cells, whose functions include detoxification of blood, production of bile to metabolize fats, and synthesis of proteins essential for metabolism, must express a specific set of genes to perform their functions. Gene expression also varies with...
16.2K
Cell Lines01:16

Cell Lines

10.0K
A cell line is a population of cells grown in vitro that can be subcultured over several generations. Normal cells cease to divide after a certain number of cell divisions, a process known as replicative senescence. This number, called the Hayflick limit, was conceptualized by Leonard Hayflick in 1961 when he observed that fetal cells grown in culture could only divide 40-60 times. This limit is due to the shortening of the telomeres during each round of cell division, preventing cell division...
10.0K
Genetic Lingo01:11

Genetic Lingo

113.7K
Overview
113.7K
Improving Translational Accuracy02:07

Improving Translational Accuracy

14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Universal cell embedding provides a foundation model for cell biology.

Nature·2026
Same author

TranscriptFormer: A generative cell atlas across 1.5 billion years of evolution.

Science (New York, N.Y.)·2026
Same author

Tabula Sapiens reveals the non-coding RNA landscape across 22 human organs and tissues.

bioRxiv : the preprint server for biology·2026
Same author

Scalable single-cell total RNA sequencing unifies coding and noncoding transcriptomics.

Nature biotechnology·2026
Same author

Spatially structured inflammatory response in the presence of a uniform stimulus.

Proceedings of the National Academy of Sciences of the United States of America·2026
Same author

Pareto optimality reveals an atlas of cellular archetypes.

Proceedings of the National Academy of Sciences of the United States of America·2026

Related Experiment Video

Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K

Benchmarking cell type and gene set annotation by large language models with AnnDictionary.

George Crowley1, , Stephen R Quake2,3,4

  • 1Department of Bioengineering, Stanford University, Stanford, California, USA.

Nature Communications
|October 29, 2025
PubMed
Summary

AnnDictionary, an open-source package, enables parallel analysis of anndata using large language models (LLMs). It benchmarks LLMs for cell-type and gene set annotation, finding over 80% accuracy for major cell types.

More Related Videos

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K

Related Experiment Videos

Last Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K

Area of Science:

  • Computational Biology
  • Bioinformatics
  • Artificial Intelligence

Background:

  • Single-cell RNA sequencing (scRNA-seq) generates large anndata datasets.
  • Automated analysis of anndata, particularly cell-type annotation, is crucial.
  • Large language models (LLMs) show potential for biological data analysis.

Purpose of the Study:

  • Introduce AnnDictionary, an open-source package for parallel anndata analysis using LLMs.
  • Benchmark major LLMs for de novo cell-type annotation accuracy.
  • Evaluate LLM performance in functional gene set annotation.

Main Methods:

  • Developed AnnDictionary, integrating LangChain and AnnData, supporting multiple LLM providers.
  • Implemented multithreading optimizations for efficient analysis of large anndata.
  • Conducted benchmarking studies comparing LLM annotations against manual annotations and gene set databases.

Main Results:

  • LLM performance in cell-type annotation varies with model size, with high agreement (>80-90%) for major cell types.
  • Inter-LLM agreement also correlates with model size.
  • Claude 3.5 Sonnet achieved high accuracy (>80%) in functional gene set annotation.

Conclusions:

  • AnnDictionary facilitates efficient, parallel analysis of anndata using LLMs.
  • LLMs demonstrate significant potential for accurate cell-type and functional annotation in scRNA-seq data.
  • Ongoing benchmarking and leaderboard maintenance will track LLM advancements in this field.