Related Experiment Video
Updated: Feb 2, 2026

10:56
A User-friendly and Powerful R Analysis of Large-scale Datasets
Published on: November 4, 2025
368
Bioinformatics Methods to Select Prognostic Biomarker Genes from Large Scale Datasets: A Review
Rémy Jardillier1,2, Florent Chatelain2, Laurent Guyon1
1University Grenoble Alpes, CEA, INSERM, Biology of Cancer Infection UMR_S 1036, 38000, Grenoble, France.
Biotechnology Journal
|November 21, 2018
Summary
Identifying reliable prognostic biomarkers from complex survival data is challenging due to a high false discovery rate. Recent advancements in variable selection, p-value definition, and biological network integration offer promising solutions for clinical validation.
Area of Science:
- Bioinformatics
- Genomics
- Biostatistics
Background:
- Survival datasets with molecular and clinical data are abundant, leading to numerous proposed prognostic biomarkers.
- Clinical validation and routine use of these biomarkers remain limited, largely attributed to a high false discovery rate.
- The high number of genes tested relative to patient cohorts contributes to this challenge.
Purpose of the Study:
- To review recent methodologies for improving prognostic biomarker discovery from survival data.
- To highlight advancements in variable selection, p-value definition, and biological knowledge integration.
- To illustrate concepts with a renal cancer dataset and provide supporting scripts.
Main Methods:
- Review of historical and recent methodologies in prognostic biomarker discovery.
- Discussion of variable selection techniques, including improved lasso penalization.
- Exploration of accurate p-value definition, false discovery rate control, and network/pathway incorporation.
Main Results:
- Recent developments focus on three key areas: enhanced variable selection, refined p-value and false discovery rate control, and integration of biological networks.
- These advancements aim to improve the reliability and clinical applicability of prognostic biomarkers.
- Illustrative examples using a renal cancer dataset are provided.
Conclusions:
- New methodologies show promise for more accurate prognostic biomarker identification.
- Independent benchmarking with diverse datasets is crucial for validating these developments.
- Further methodological research is necessary to advance the field.
Related Concept Videos
Review and Preview
8.4K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
8.4K
Review and Preview
11.3K
Data are individual items of information obtained from a population or sample. Data may be classified as qualitative (categorical), quantitative continuous, or quantitative discrete. Because it is not practical to measure the entire population in a study, researchers use samples to represent the population. A random sample is a representative group from the population chosen by using a method that gives each individual in the population an equal chance of being included in the sample. Random...
11.3K
pH Scale
79.7K
Hydronium and hydroxide ions are present both in pure water and in all aqueous solutions, and their concentrations are inversely proportional as determined by the ion product of water (Kw). The concentrations of these ions in a solution are often critical determinants of the solution’s properties and the chemical behaviors of its other solutes. Two different solutions can differ in their hydronium or hydroxide ion concentrations by a million, billion, or even trillion times. A common means of...
79.7K
Antibiotic Selection
59.9K
Overview
59.9K
Scaling
593
In designing and analyzing filters, resonant circuits, or circuit analysis at large, working with standard element values like 1 ohm, 1 henry, or 1 farad can be convenient before scaling these values to more realistic figures. This approach is widely utilized by not employing realistic element values in numerous examples and problems; it simplifies mastering circuit analysis through convenient component values. The complexity of calculations is thereby reduced, with the understanding that...
593
What is Natural Selection?
129.1K
Natural selection is an evolutionary process in which individuals with survival-promoting traits reproduce at higher rates. These favorable traits become more common within a population or species. Naturally selected traits initially arise via random genetic mutations. In order for selection to occur, there must be variation within a population, the trait controlling the variation must be heritable, and there must be an evolutionary advantage for variation in the trait.
129.1K

