Related Experiment Video
Updated: Sep 3, 2025

Micromanipulation of Circulating Tumor Cells for Downstream Molecular Analysis and Metastatic Potential Assessment
Published on: May 14, 2019
Visual Clustering of Transcriptomic Data from Primary and Metastatic Tumors-Dependencies and Novel Pitfalls
André Marquardt1,2,3, Philip Kollmannsberger4, Markus Krebs5,6
1Institute of Pathology, Klinikum Stuttgart, 70174 Stuttgart, Germany.
Abstract:
Personalized oncology is a rapidly evolving area and offers cancer patients therapy options that are more specific than ever. However, there is still a lack of understanding regarding transcriptomic similarities or differences of metastases and corresponding primary sites. Applying two unsupervised dimension reduction methods (t-Distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP)) on three datasets of metastases (n =682 samples) with three different data transformations (unprocessed, log10 as well as log10 + 1 transformed values), we visualized potential underlying clusters. Additionally, we analyzed two datasets (n =616 samples) containing metastases and primary tumors of one entity, to point out potential familiarities. Using these methods, no tight link between the site of resection and cluster formation outcome could be demonstrated, or for datasets consisting of solely metastasis or mixed datasets. Instead, dimension reduction methods and data transformation significantly impacted visual clustering results. Our findings strongly suggest data transformation to be considered as another key element in the interpretation of visual clustering approaches along with initialization and different parameters. Furthermore, the results highlight the need for a more thorough examination of parameters used in the analysis of clusters.
Insights
Data transformation significantly impacts transcriptomic clustering in cancer. Understanding these transformations is crucial for interpreting results in personalized oncology and metastasis research.
Area of Science:
- Genomics
- Bioinformatics
- Cancer Research
Background:
- Personalized oncology aims for targeted cancer therapies.
- Understanding transcriptomic differences between primary tumors and metastases is limited.
- Current research lacks clarity on how data transformations affect clustering analyses.
Purpose of the Study:
- To investigate transcriptomic similarities and differences between metastases and primary tumors.
- To evaluate the impact of dimension reduction techniques and data transformations on clustering results.
- To identify key factors influencing the interpretation of transcriptomic data in cancer research.
Main Methods:
- Applied t-Distributed Stochastic Neighbor Embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP) for dimension reduction.
- Utilized three datasets of metastases (n=682) and two datasets of primary tumors and metastases (n=616).
- Analyzed unprocessed, log10, and log10 + 1 transformed data values to assess transformation effects.
Main Results:
- No significant link was found between resection site and cluster formation.
- Dimension reduction methods and data transformation methods significantly influenced visual clustering outcomes.
- The choice of data transformation critically affected the observed clustering patterns.
Conclusions:
- Data transformation is a critical factor in interpreting visual clustering of transcriptomic data.
- Initialization and parameter choices also significantly impact clustering results.
- Further investigation into parameters for cluster analysis is necessary for robust findings in cancer research.

