Related Experiment Videos
Sparse Logistic Regression on Genomic Data for Prediction of Tumour Pathological Subtype
Abstract:
The correct prediction of tumour subtype is critical for the treatment of cancer patients to maximise the chance of survival. The patients' genomic information, such as copy number alterations (CNA) profile, has increasingly become an important factor in the prediction to supplement the traditional pathological subtyping. The incorporation of the CNA information in a prediction model, such as logistic regression, faces two major statistical challenges: first, how to estimate the model parameters in the thousands and, second, how to deal with the correlation of CNA between genomic regions. To address them, we propose a sparse logistic regression model with random effects where some of its parameters are estimated to zero while the other parameters are non-zero. In effect, a variable selection is embedded in the modelling. To deal with the correlation of CNA across genomic regions, we extend further the model to incorporate an additional penalty in the corresponding likelihood function in the logistic regression. The results show that we can identify selected genomic regions that are informative to distinguish different tumour subtypes, while giving a good prediction ability. We illustrate the methodology using CNA dataset from a lung cancer cohort.