Biomarker discovery and redundancy reduction towards classification using a multi-factorial MALDI-TOF MS T2DM mouse
Chris Bauer1, Frank Kleinjung, Celia J Smith
1MicroDiscovery GmbH, Marienburger Str, 1, 10405 Berlin, Germany. chris.bauer@microdiscovery.de
BMC Bioinformatics
|May 11, 2011
Summary
This study introduces a bioinformatics workflow to analyze complex diabetes data, identifying specific peptide biomarkers linked to genotype and diet interactions. The method offers efficient feature selection for improved classification.
Area of Science:
- Bioinformatics
- Proteomics
- Systems Biology
Background:
- Diabetes mellitus is a complex, multifactorial disease requiring sophisticated analytical approaches.
- Multi-factorial studies generate complex datasets with inherent redundancy, posing significant bioinformatics challenges.
- Analyzing high-dimensional data, such as from proteomic studies, necessitates advanced computational methods.
Purpose of the Study:
- To develop and present a comprehensive bioinformatics workflow for analyzing complex, multi-factorial experimental data.
- To identify specific molecular markers associated with distinct combinations of experimental factors, such as genotype and diet.
- To leverage data redundancy for effective feature selection and classification in biological studies.
Main Methods:
- Development of a novel bioinformatics workflow integrating analysis of variance (ANOVA) and redundancy exploitation.
- Application of the workflow to analyze proteomic data from a polygenic mouse model of diet-induced type 2 diabetes.
- Utilizing peptide correlation for feature selection and classification model development.
Main Results:
- The workflow successfully identified peptides with significant fold changes specific to the interaction of particular mouse strains and diets.
- Exploitation of redundancy facilitated intuitive visualization of peptide correlations and natural feature selection.
- Classification models built on the selected features demonstrated performance comparable to methods using more complex feature selection strategies.
Conclusions:
- The combined use of ANOVA and redundancy exploitation enables robust identification of biomarker candidates in complex, multi-dimensional mass spectrometry profiling studies.
- The proposed feature selection method offers a computationally efficient and intuitive alternative to global optimization strategies with similar predictive performance.
- The developed workflow is implemented in R, with scripts available for broader research application.
