Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Modern Molecular Taxonomy01:29

Modern Molecular Taxonomy

Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
RNA-seq03:21

RNA-seq

RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Respiratory virus genomic epidemiology during post-pandemic re-emergence of influenza in Australia.

Virus evolution·2026
Same author

Intelligent Electrochemical Sensing: Machine Learning-Powered Multidimensional Fingerprinting for Simultaneous Detection of Six Antibiotics in Complex Matrices.

Analytical chemistry·2026
Same author

Automated dairy cattle body condition score using side-view images and deep learning.

Journal of dairy science·2026
Same author

Clec3b⁺ fibroblasts are the primary effectors of portal fibrosis following activation via a KLF4/periostin axis.

Nature communications·2026
Same author

Investigating the Impact of Host Genetics on the Risk of Disease Progression in Individuals With Influenza.

Immunity, inflammation and disease·2026
Same author

The E3 ubiquitin ligase NEDD4 protects against nonesterified fatty acid-induced hepatic inflammatory injury in ketotic cows.

Journal of dairy science·2026

Related Experiment Video

Updated: May 22, 2026

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

Linear normalised hash function for clustering gene sequences and identifying reference sequences from multiple

Manal Helal1, Fanrong Kong, Sharon Ca Chen

  • 1Sydney Emerging Infections and Biosecurity Institute, Sydney Medical School - Westmead, University of Sydney, Sydney, New South Wales, Australia. vitali.sintchenko@swahs.health.nsw.gov.au.

Microbial Informatics and Experimentation
|May 17, 2012
PubMed
Summary

A new method using linear mapping hash function and multiple sequence alignment (MSA) efficiently clusters gene sequences. This approach accurately identifies optimal clusters and centroids for diverse genomic datasets, improving classification in comparative genomics.

More Related Videos

A Practical Guide to Phylogenetics for Nonexperts
12:00

A Practical Guide to Phylogenetics for Nonexperts

Published on: February 5, 2014

Related Experiment Videos

Last Updated: May 22, 2026

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

A Practical Guide to Phylogenetics for Nonexperts
12:00

A Practical Guide to Phylogenetics for Nonexperts

Published on: February 5, 2014

Area of Science:

  • Bioinformatics
  • Genomics
  • Computational Biology

Background:

  • Comparative genomics requires robust sequence similarity assessment and clustering for classification.
  • Challenges exist in defining optimal cluster numbers, density, and boundaries for polymorphic gene sequences.
  • Existing methods struggle with variable sequence datasets and determining optimal cluster parameters.

Purpose of the Study:

  • To develop a novel method for identifying cluster centroids and the optimal number of clusters.
  • To create a universally applicable method for diverse sequence datasets.
  • To address limitations in current gene sequence clustering techniques.

Main Methods:

  • Developed a novel method combining a linear mapping hash function with multiple sequence alignment (MSA).
  • Utilized MSA to sort sequences by similarity, enabling identification of cluster cut-offs and centroids.
  • Employed the linear mapping hash function to analyze distance matrices and identify optimal clustering boundaries.

Main Results:

  • The method successfully identified optimal cluster numbers, cut-offs, and centroids for both closely related and highly variable sequences.
  • Outperformed existing unsupervised machine learning and dimensionality reduction methods in clustering accuracy.
  • Demonstrated scalability, handling clusters of various sizes and shapes without prior knowledge of cluster count or distance.

Conclusions:

  • The combined MSA and linear mapping hash function offers a computationally efficient gene sequence clustering solution.
  • This method is valuable for assessing similarity, clustering microbial genomes, and identifying reference sequences.
  • It provides a powerful tool for evolutionary studies of bacteria and viruses.