Related Experiment Video
Updated: Oct 14, 2025

05:22
Analyzing Multifactorial RNA-Seq Experiments with DiCoExpress
Published on: July 29, 2022
3.7K
Band-based similarity indices for gene expression classification and clustering
1Departamento de Matemáticas, Instituto Gregorio Millán, Universidad Carlos III de Madrid, 28911, Leganés, Spain. etorrent@est-econ.uc3m.es.
Scientific Reports
|November 4, 2021
Summary
Modified Band Depth (MBD) offers a computationally efficient way to measure similarity in high-dimensional data. This novel approach, utilizing data bands, outperforms traditional methods like Euclidean distance in classification and clustering tasks.
Area of Science:
- Multivariate Data Analysis
- Computational Statistics
- Bioinformatics
Background:
- Traditional depth measures for multivariate data are computationally infeasible in high dimensions.
- Modified Band Depth (MBD) is an exception, proving effective for high-dimensional gene expression data analysis.
- Assessing data point centrality is crucial for understanding multivariate data structures.
Purpose of the Study:
- To develop and evaluate novel band-based similarity indices for multivariate data.
- To assess the computational efficiency and performance of these indices in high-dimensional settings.
- To compare the effectiveness of band-based similarity with classical distance metrics.
Main Methods:
- Utilized Modified Band Depth (MBD) to define centrality and relationships within data bands.
- Constructed binary matrices and contingency tables from data bands to quantify pairwise (dis)similarity.
- Derived standard similarity indices from contingency tables for analysis.
Main Results:
- The proposed band-based similarity approach is computationally efficient and scalable to high dimensions.
- Several derived similarity indices demonstrated superior performance compared to classical distances, including Euclidean distance.
- The method showed strong results in classification and clustering tasks across simulated and real datasets.
Conclusions:
- Band-based similarity indices derived from Modified Band Depth offer a powerful alternative for high-dimensional data analysis.
- The technique provides a computationally efficient and effective tool for uncovering patterns in complex datasets.
- The framework is flexible and extensible to other similarity coefficients and applications beyond classification/clustering.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
6.4K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.4K
DNA Microarrays
19.0K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
19.0K
Modern Molecular Taxonomy
262
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
262
Cell Specific Gene Expression
5.0K
5.0K
RNA-seq
10.5K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.5K

