Related Experiment Video
Updated: Jun 1, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A comparison of methods for data-driven cancer outlier discovery, and an application scheme to semisupervised
Seppo Karrila1, Julian Hock Ean Lee, Greg Tucker-Kellogg
1Lilly Singapore Centre for Drug Discovery, Eli Lilly and Company, Singapore.
Abstract:
A core component in translational cancer research is biomarker discovery using gene expression profiling for clinical tumors. This is often based on cell line experiments; one population is sampled for inference in another. We disclose a semisupervised workflow focusing on binary (switch-like, bimodal) informative genes that are likely cancer relevant, to mitigate this non-statistical problem. Outlier detection is a key enabling technology of the workflow, and aids in identifying the focus genes.We compare outlier detection techniques MOST, LSOSS, COPA, ORT, OS, and t-test, using a publicly available NSCLC dataset. Removing genes with Gaussian distribution is computationally efficient and matches MOST particularly well, while also COPA and OS pick prognostically relevant genes in their top ranks. Also our stability assessment is in favour of both MOST and COPA; the latter does not pair well with prefiltering for non-Gaussianity, but can handle data sets lacking non-cancer cases.We provide R code for replicating our approach or extending it.
Related Concept Videos
Cancer Survival Analysis
Kaplan-Meier Approach
Comparing the Survival Analysis of Two or More Groups
Quantifying and Rejecting Outliers: The Grubbs Test
