Related Experiment Video
Updated: May 22, 2026

12:27
Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Interpolative multidimensional scaling techniques for the identification of clusters in very large sequence sets
Adam Hughes1, Yang Ruan, Saliya Ekanayake
1Pervasive Technology Institute, Indiana University, Bloomington, IN 47408, USA.
BMC Bioinformatics
|April 28, 2012
Summary
This study introduces an efficient pipeline for clustering large bacterial 16S rRNA sequence sets. Using interpolative multidimensional scaling (MDS) significantly reduces computational time for accurate gene cluster identification.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Modern pyrosequencing enables direct analysis of complex bacterial populations, like 16S rRNA, from various samples.
- Analyzing large datasets (100,000+ sequences) for gene clusters is computationally challenging.
- Efficient pipelines are needed for clustering vast sequence read sets.
Purpose of the Study:
- To develop an efficient computational pipeline for clustering large sequence read sets.
- To identify potential gene clusters and families within massive biosequence data.
Main Methods:
- Utilized pairwise alignment techniques (Needleman-Wunsch) to calculate genetic distances.
- Implemented novel interpolative multidimensional scaling (MDS) for data visualization and clustering.
- Compared interpolative MDS with traditional multiple sequence alignment (MSA) and full MDS.
Main Results:
- Interpolative MDS achieved qualitatively similar clustering results to full MDS.
- Reduced the time to cluster 100,000 sequences from seven hours to under one hour.
- Demonstrated substantial computational cost savings for large-scale sequence analysis.
Conclusions:
- Interpolative MDS offers significant computational advantages for clustering large sequence datasets.
- Further optimization of training set size for interpolative MDS is warranted.
- This approach enables future clustering of even larger sequence sets efficiently.
Related Concept Videos
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Multi-species Conserved Sequences
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
