Related Experiment Video
Updated: Jun 17, 2026

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
Modelling p-value distributions to improve theme-driven survival analysis of cancer transcriptome datasets
Esteban Czwan1, Benedikt Brors, David Kipling
1School of Medicine, Cardiff University, Heath Park, Cardiff CF144XN, UK.
Theme-driven cancer survival studies often yield false positives due to incorrect assumptions about gene set p-values. This study introduces a new method to accurately assess gene set prognostic power, identifying novel biological themes linked to patient survival.
Area of Science:
- Bioinformatics
- Cancer Genomics
- Translational Oncology
Background:
- Theme-driven cancer survival studies aim to link gene expression signatures to patient prognosis.
- Current methods often fail to adequately test the biological relevance of gene sets, potentially leading to false conclusions.
- An incorrect assumption of uniform p-value distributions for random gene sets is a key limitation.
Purpose of the Study:
- To develop and validate a robust method for theme-driven cancer survival analysis.
- To address the pitfall of non-uniform p-value distributions in assessing gene set prognostic power.
- To identify novel biological themes and pathways with significant prognostic value in cancer.
Main Methods:
- An automated theme-driven method was developed using a permutation approach to empirically approximate p-value distributions.
- The method assesses predefined biologically-related gene sets for association with patient survival.
- Comparative analysis with existing methods highlighted issues with false positive rates in prior studies.
Main Results:
- Non-uniform p-value distributions were confirmed, indicating a significant problem with false positive rates in previous studies.
- Novel ontological categories with prognostic power were identified in two public cancer datasets.
- Specific findings include associations between "fatty acid metabolism" and breast cancer survival, and "receptor mediated endocytosis", "brain development", "apical plasma membrane", and "MAPK signaling pathway" with lung cancer survival.
Conclusions:
- The assumption of uniform p-values for random gene sets in current survival studies is flawed and can lead to erroneous results.
- The developed approach corrects for this pitfall, offering a more reliable method for identifying prognostic gene sets.
- This work opens new avenues for discovering higher-level biological themes and pathways with prognostic implications in clinical data.
Related Concept Videos
Cancer Survival Analysis
Comparing the Survival Analysis of Two or More Groups
Assumptions of Survival Analysis
Kaplan-Meier Approach
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...