Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genome Annotation and Assembly03:36

Genome Annotation and Assembly

18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Leaky Scanning02:28

Leaky Scanning

5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA.  Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Improving Translational Accuracy02:07

Improving Translational Accuracy

10.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.2K
RNA-seq03:21

RNA-seq

9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Viral Mutations00:36

Viral Mutations

32.3K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
32.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Multisite Assessment of Methods for Cell Preservation Upstream of Single-Cell RNA Sequencing.

Journal of biomolecular techniques : JBT·2026
Same author

Single amino acid substitution in DNA Polymerase I dramatically alters infection dynamics of bacteriophage T7.

bioRxiv : the preprint server for biology·2026
Same author

Anellovirus-Mediated Interferon Dysregulation Enhances Virus-Induced Lung Injury.

American journal of respiratory cell and molecular biology·2026
Same author

A short-term, randomized, controlled, feasibility study of the effects of different vegetables on the gut microbiota and microRNA expression in infants.

Frontiers in microbiomes·2026
Same author

Influence of Singular First Foods on the Infant Gut Microbiome: A Randomized Controlled Trial.

The Journal of nutrition·2026
Same author

An Integrative Proteotranscriptomics Approach Reveals New ADAM9 Substrates and Downstream Pathways.

Molecular & cellular proteomics : MCP·2026

Related Experiment Video

Updated: Jun 27, 2025

Validating Whole Genome Nanopore Sequencing, using Usutu Virus as an Example
05:45

Validating Whole Genome Nanopore Sequencing, using Usutu Virus as an Example

Published on: March 11, 2020

8.8K

Improvements in viral gene annotation using large language models and soft alignments.

William L Harrigan1, Barbra D Ferrell2, K Eric Wommack2

  • 1Hawai'i Institute of Marine Biology, University of Hawai'i at Mānoa, Honolulu, HI, 96822, USA.

BMC Bioinformatics
|April 25, 2024
PubMed
Summary

Large Language Models (LLMs) offer a new way to annotate viral protein sequences. This embedding-based approach improves accuracy and interpretability over traditional methods.

Keywords:
AlignmentsLarge language modelsProtein homologyViruses

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

548
Phage Phenomics: Physiological Approaches to Characterize Novel Viral Proteins
09:40

Phage Phenomics: Physiological Approaches to Characterize Novel Viral Proteins

Published on: June 11, 2015

12.2K

Related Experiment Videos

Last Updated: Jun 27, 2025

Validating Whole Genome Nanopore Sequencing, using Usutu Virus as an Example
05:45

Validating Whole Genome Nanopore Sequencing, using Usutu Virus as an Example

Published on: March 11, 2020

8.8K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

548
Phage Phenomics: Physiological Approaches to Characterize Novel Viral Proteins
09:40

Phage Phenomics: Physiological Approaches to Characterize Novel Viral Proteins

Published on: June 11, 2015

12.2K

Area of Science:

  • Molecular Biology
  • Bioinformatics
  • Artificial Intelligence

Background:

  • Protein sequence annotation is a significant challenge in molecular biology, especially for viral proteins.
  • Traditional homology search methods (alignment, k-mer, profile-based) struggle with viral sequences due to limited homology.
  • Large Language Models (LLMs) present a novel approach for protein sequence annotation using embeddings.

Purpose of the Study:

  • To introduce a novel methodology for protein sequence annotation using LLM embeddings.
  • To address the limitations of traditional methods in annotating viral proteins.
  • To enhance the efficiency and interpretability of protein annotation.

Main Methods:

  • Development of a soft alignment algorithm leveraging amino acid embedding similarity.
  • Bypassing traditional scoring matrices by using embedding similarity.
  • Creating transparent, BLAST-like alignment visualizations for interpretability.

Main Results:

  • The soft alignment algorithm surpasses pooled embedding-based models in efficiency and interpretability.
  • The method allows users to trace homologous amino acids and provides clear alignment visualizations.
  • The novel approach successfully annotated sequences that blastp and pooling-based methods failed to detect, as validated on Virus Orthologous Groups and ViralZone databases.

Conclusions:

  • LLM-based embedding approaches hold significant potential for improving protein sequence annotation, particularly in viral genomics.
  • This method offers a more efficient and accurate pathway for protein function inference in molecular biology.
  • The findings highlight a promising advancement in combining AI with traditional biological research for enhanced annotation.