Related Experiment Video
Updated: Jun 5, 2026

JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
Clustering gene expression data with a penalized graph-based metric
Ariel E Bayá1, Pablo M Granitto
1CIFASIS French Argentine International Center for Information and Systems Sciences, UPCAM (France)/UNR-CONICET (Argentina), Bv 27 de Febrero 210 Bis, 2000 Rosario, República Argentina. baya@cifasis-conicet.gov.ar
This study introduces a new Penalized k-Nearest-Neighbor-Graph (PKNNG) metric for improved clustering of complex, high-dimensional "-omic" data. The PKNNG metric enhances existing clustering algorithms, offering accurate and interpretable results for gene expression datasets.
Area of Science:
- Bioinformatics
- Computational Biology
- Data Science
Background:
- Clustering microarray datasets is crucial for -omic sciences.
- Handling manifold structures in high-dimensional data presents a significant clustering challenge.
- Gene expression datasets often exhibit complex, non-compact data structures.
Purpose of the Study:
- Introduce a novel distance metric for clustering high-dimensional data with manifold structures.
- Enhance the performance of existing clustering algorithms for -omic data analysis.
- Provide a user-friendly tool for accurate and interpretable data clustering.
Main Methods:
- Developed the Penalized k-Nearest-Neighbor-Graph (PKNNG) based metric.
- PKNNG metric employs a two-step graph construction with penalized edge weights.
- Evaluated the metric's performance on public gene expression datasets and simulated data.
Main Results:
- The PKNNG metric demonstrated promising clustering results across various datasets.
- It effectively handles data with manifold structures, improving clustering accuracy.
- PKNNG metric-based clustering achieved performance comparable to advanced algorithms.
Conclusions:
- The PKNNG metric offers a significant improvement for pairwise-distance based clustering methods.
- Researchers can integrate PKNNG with existing algorithms like hierarchical clustering for better dendrograms.
- This approach enhances interpretability and accuracy in high-dimensional data analysis without requiring new methods.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Quantifying and Rejecting Outliers: The Grubbs Test
