A consensus orthogonal partial least squares discriminant analysis (OPLS-DA) strategy for multiblock Omics data
Julien Boccard1, Douglas N Rutledge
1Laboratoire de Chimie Analytique, AgroParisTech, Paris, France.
Analytica Chimica Acta
|March 19, 2013
Summary
This study introduces a new method for fusing multiple omics data types, improving biological system analysis. The approach combines multiple kernel learning and orthogonal partial least squares discriminant analysis for efficient data integration.
Area of Science:
- Systems Biology
- Bioinformatics
- Analytical Chemistry
Background:
- Omics approaches offer broad biological system monitoring but no single technique captures full biochemical content.
- Complex biological data require fusion of information from multiple sources for comprehensive analysis.
- High-dimensional omics data present challenges in knowledge extraction.
Purpose of the Study:
- To develop a generic methodology for fusing omics data from multiple sources.
- To provide an efficient tool for handling high-dimensional and noisy omics datasets.
- To assess the potential of the proposed fusion method through real-world case studies.
Main Methods:
- A novel methodology combining multiple kernel learning and Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA).
- Application of the fusion method to diverse omics datasets, including metabolomics, transcriptomics, and proteomics.
- Validation using three case studies: plant metabolomics, grape variety classification, and cancer cell line data analysis.
Main Results:
- Demonstrated successful fusion of mass spectrometry-based metabolomic data from different ionization modes.
- Successfully classified wine grape varieties using 2D heteronuclear magnetic resonance spectroscopy data.
- Effectively combined heterogeneous systems biology data, including publicly available NCI-60 cancer cell line omics data.
Conclusions:
- The proposed method is a relevant and widely applicable tool for efficient fusion of multiple omics data.
- Integrating omics data from different sources provides a more complete view of biological systems.
- The method effectively handles the inherent characteristics of multiple omics data, such as noisy collinear variables.

