Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Genetic Lingo01:11

Genetic Lingo

104.7K
Overview
104.7K
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

14.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
14.2K
Genome Annotation and Assembly03:36

Genome Annotation and Assembly

19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
Language Development01:22

Language Development

450
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
450
Genomics02:02

Genomics

37.5K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
37.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Discovery of BilV reveals a multienzymatic basis for bilirubin reduction across vertebrate gut microbiomes.

bioRxiv : the preprint server for biology·2026
Same author

Isolation of a highly virulent colibactin-positive tumor-promoting strain of <i>Escherichia coli</i> from the gut microbiota of an adult.

mSphere·2026
Same author

SpiR is a gut microbial enzyme that drives cholesterol conversion.

Nature communications·2026
Same author

LAMBDA: A Prophage Detection Benchmark for Genomic Language Models.

bioRxiv : the preprint server for biology·2026
Same author

Differences in antibiotic treatment for children hospitalized with pneumonia.

Journal of hospital medicine·2026
Same author

Functional analyses of bacterial NanoRNase B proteins reveals defining features of this enzyme family.

Nucleic acids research·2025

Related Experiment Video

Updated: Sep 12, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

900

The Impact of Tokenizer Selection in Genomic Language Models.

LeAnn M Lindsey1,2, Nicole L Pershing3, Anisa Habib1

  • 1Kahlert School of Computing, University of Utah, SLC, UT, USA.

Biorxiv : the Preprint Server for Biology
|August 8, 2025
PubMed
Summary

Character tokenization is best for nucleotide-level genomic tasks, outperforming sub-word methods. Byte-pair encoding showed promise for specific applications like SARS-CoV-2 variant classification.

Keywords:
genomic language modelsgenomicstokenization

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

682
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

579

Related Experiment Videos

Last Updated: Sep 12, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

900
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

682
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

579

Area of Science:

  • Computational biology
  • Genomic data analysis
  • Artificial intelligence in genomics

Background:

  • Genomic language models (GLMs) are emerging tools for genetic sequence analysis.
  • GLMs use various tokenization methods, including character, k-mer, and byte-pair encoding (BPE).
  • Genomic sequences present unique challenges compared to natural language, impacting tokenization strategies.

Purpose of the Study:

  • To investigate the impact of different tokenization methods on GLM performance.
  • To evaluate tokenization strategies across diverse genomic classification tasks.
  • To compare character tokenization with sub-word tokenization (BPE) in the Mamba model.

Main Methods:

  • Evaluated downstream performance of GLMs on 44 classification fine-tuning tasks.
  • Conducted a direct comparison between byte-pair encoding and character tokenization within the Mamba state-space model.
  • Assessed model performance based on task-specific metrics.

Main Results:

  • Character tokenization outperformed sub-word methods on tasks requiring nucleotide-level resolution, such as splice site prediction and promoter detection.
  • Byte-pair encoding demonstrated superior performance in SARS-CoV-2 variant classification.
  • Limited statistically significant differences were observed between tokenization methods for most other downstream tasks.

Conclusions:

  • The choice of tokenization significantly impacts GLM performance, particularly for nucleotide-level genomic tasks.
  • Character tokenization is a robust choice for tasks demanding high resolution.
  • Further research is needed to optimize tokenization for diverse genomic applications and model architectures.