Goistrat: gene-of-interest-based sample stratification for the evaluation of functional differences
Carlos Uziel Pérez Malla1,2, Jessica Kalla1, Andreas Tiefenbacher1,3
1Department of Pathology, Medical University of Vienna, Währinger Gürtel 18-20, Vienna, 1090, Austria.
Purpose:
Understanding the impact of gene expression in pathological processes, such as carcinogenesis, is crucial for understanding the biology of cancer and advancing personalised medicine. Yet, current methods lack biologically-informed-omics approaches to stratify cancer patients effectively, limiting our ability to dissect the underlying molecular mechanisms.
Results:
To address this gap, we present a novel workflow for the stratification and further analysis of multi-omics samples with matched RNA-Seq data that relies on MSigDB curated gene sets, graph machine learning and ensemble clustering. We compared the performance of our workflow in the top 8 TCGA datasets and showed its clear superiority in separating samples for the study of biological differences. We also applied our workflow to analyse nearly a thousand prostate cancer samples, focusing on the varying expression of the FOLH1 gene, and identified specific pathways such as the PI3K-AKT-mTOR gene sets as well as signatures linked to prostate tumour aggressiveness.
Conclusion:
Our comprehensive approach provides a novel tool to identify disease-relevant functions of genes of interest (GOI) in large datasets. This integrated approach offers a valuable framework for understanding the role of the expression variation of a GOI in complex diseases and for informing on targeted therapeutic strategies.
Insights
This study introduces a new multi-omics analysis workflow using graph machine learning to effectively stratify cancer patients. The method improves the study of biological differences and identifies key pathways in prostate cancer progression.
Area of Science:
- Genomics and Bioinformatics
- Computational Biology
- Cancer Research
Background:
- Understanding gene expression in cancer is vital for personalized medicine.
- Current omics methods struggle to effectively stratify cancer patients for molecular analysis.
Purpose of the Study:
- To develop a novel, biologically-informed workflow for multi-omics data stratification.
- To improve the analysis of cancer patient samples and identify disease mechanisms.
Main Methods:
- Utilized a workflow combining MSigDB gene sets, graph machine learning, and ensemble clustering.
- Applied the method to multi-omics samples with matched RNA-Seq data.
- Validated performance on eight TCGA datasets and analyzed nearly a thousand prostate cancer samples.
Main Results:
- Demonstrated superior sample separation for studying biological differences compared to existing methods.
- Identified specific pathways, including PI3K-AKT-mTOR gene sets, linked to prostate tumor aggressiveness.
- Successfully analyzed the impact of FOLH1 gene expression variation.
Conclusions:
- The developed workflow is a novel tool for identifying disease-relevant gene functions in large datasets.
- Provides a framework for understanding gene expression variation in complex diseases.
- Informs the development of targeted therapeutic strategies.
Related Concept Videos
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...


