Classification based on extensions of LS-PLS using logistic regression: application to clinical and multiple genomic
Caroline Bazzoli1, Sophie Lambert-Lacroix2
1Laboratoire Jean Kuntzman, Univ. Grenoble-Alpes, 700 avenue centrale, Saint Martin d'Hères, 38401, France. caroline.bazzoli@univ-grenoble-alpes.fr.
BMC Bioinformatics
|September 8, 2018
Summary
Combining clinical and genomic data improves prediction accuracy for high-dimensional datasets. New methods using least squares-partial least squares (LS-PLS) demonstrate stable and encouraging results in cancer studies.
Area of Science:
- Bioinformatics
- Genomic Data Analysis
- Biostatistics
Background:
- High-dimensional genomic data analysis often excludes valuable clinical information.
- Integrating clinical and genomic data is suggested to enhance predictive accuracy.
- Existing prediction methods frequently analyze genomic or clinical data in isolation.
Purpose of the Study:
- To develop and evaluate classification methods that simultaneously utilize both clinical and genomic data.
- To apply dimensionality reduction techniques specifically to high-dimensional genomic variables.
- To compare the predictive performance of integrated approaches against single-data-type methods.
Main Methods:
- Proposed one-step classification approaches based on extensions of the least squares-partial least squares (LS-PLS) method for logistic regression.
- Utilized partial least squares (PLS) for dimensionality reduction of genomic data.
- Compared prediction performances through simulations and analysis of real cancer study datasets.
Main Results:
- Methods relying solely on clinical or genomic data generally showed poor performance.
- The proposed LS-PLS methods demonstrated advantages for classification tasks.
- Prediction results were encouraging and stable across different datasets and feature selections.
Conclusions:
- Simultaneous analysis of clinical and genomic data, particularly with LS-PLS extensions, offers improved predictive power.
- The developed methods provide stable and reliable predictions for classification tasks.
- The R package lsplsGlm facilitates the application of these advanced classification techniques.
Related Concept Videos
Multiple Regression
4.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.0K
Regression Toward the Mean
7.1K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.1K
Genome Size and the Evolution of New Genes
9.1K
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
9.1K
Genomics
40.7K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
40.7K
Clinical Applications of Epidermal Stem Cells
3.3K
Epidermal stem cells (EpiSCs) are mainly located at the basal layer of the epidermis. These cells repair minor injuries of the skin and replace dead skin cells. However, EpiSCs’ cannot heal severe wounds such as major burns or those from diabetes or hereditary disorders. In such cases, culturing the epidermal stem cells from the patient is possible and has yielded successful treatment options, such as laboratory-grown skin grafts. These grafts are synthesized using a patient’s own...
3.3K
Statistical Software for Data Analysis and Clinical Trials
1.6K
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.6K


