Related Experiment Video
Updated: Jul 17, 2026

07:28
JUMPn: A Streamlined Application for Protein Co-Expression Clustering and Network Analysis in Proteomics
Published on: October 19, 2021
Assessment of hierarchical clustering methodologies for proteomic data mining
Bruno Meunier1, Emilie Dumas, Isabelle Piec
1UR 1213, Unité de Recherches sur les Herbivores, Equipe Croissance et Métabolisme du Muscle, INRA de Clermont-Ferrand/Theix, F-63122 [corrected] Saint-Genès Champanelle, France. bruno.meunier@clermont.inra.fr
Journal of Proteome Research
|January 6, 2007
Summary
Hierarchical clustering is key for exploring proteomic data. Combining Pearson correlation with Ward
Area of Science:
- Proteomics
- Bioinformatics
- Data Mining
Background:
- Hierarchical clustering is a valuable tool for initial exploration of proteomic data, grouping samples or proteins based on expression profiles.
- Clustering outcomes are sensitive to parameter choices, including data preprocessing, similarity measures, and dendrogram construction.
- Evaluating different clustering strategies is crucial for optimizing the analysis of complex biological datasets.
Purpose of the Study:
- To assess and identify optimal hierarchical clustering strategies for proteomic data analysis.
- To compare the performance of various parameter combinations using a robust quality metric.
- To provide guidance on selecting effective clustering methods for proteomic expression profiling.
Main Methods:
- Utilized hierarchical clustering methodology for proteomic data exploration.
- Employed the F-measure as a primary metric for evaluating clustering quality.
- Investigated various parameter combinations, including data preprocessing (log transformation), similarity measures (Pearson correlation), and aggregation methods (Ward's method).
Main Results:
- The combination of log-transformed data, Pearson correlation for similarity, and Ward's method for aggregation emerged as a highly effective clustering strategy.
- This specific strategy demonstrated superior performance across the studied proteomic datasets.
- PermutMatrix software, adapted from transcriptomics, was used to facilitate these analyses.
Conclusions:
- The optimal hierarchical clustering strategy for the analyzed proteomic datasets involves using log-transformed data, Pearson correlation, and Ward's method.
- This validated approach enhances the reliability of grouping samples and proteins based on expression patterns.
- The findings offer practical recommendations for researchers utilizing clustering in proteomic data mining.
Related Concept Videos
Proteomics
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Subcellular Fractionation
The homogenate obtained after cell lysis contains various membrane-bound organelles that can be further separated into pure fractions by subcellular fractionation. These isolates are used to study specific cellular components, analyze localized protein activity, and are even employed in diagnostics. Fractionation is typically achieved using centrifugation methods, the most common being density-gradient and differential centrifugation.
Differential Centrifugation
Differential centrifugation is...
Differential Centrifugation
Differential centrifugation is...
