Related Experiment Videos
Scaling and normalization effects in NMR spectroscopic metabonomic data sets
Andrew Craig1, Olivier Cloarec, Elaine Holmes
1Biological Chemistry, Faculty of Natural Sciences, Imperial College London, Sir Alexander Fleming Building, South Kensington, London SW7 2AZ U.K.
Analytical Chemistry
|April 4, 2006
Summary
Data preprocessing in metabonomics, such as normalization and scaling, is crucial but context-dependent. There is no single optimal method for all 1H NMR spectroscopic data sets from urine.
Area of Science:
- Metabonomics
- Nuclear Magnetic Resonance (NMR) Spectroscopy
- Data Science
Background:
- Metabonomics research often faces confusion regarding the necessity and methods of spectroscopic data preprocessing.
- Data in metabonomics are typically structured as tables of sample measurements, such as spectral peak intensities or metabolite concentrations.
Purpose of the Study:
- To clarify the roles of normalization (row operations) and scaling (column operations) in metabonomics data preprocessing.
- To highlight the specific challenges of analyzing urine samples using 1H NMR spectroscopy.
- To discuss the implications of data preprocessing on statistical analyses, like correlation coefficients.
Main Methods:
- Defined and discussed normalization and scaling operations for metabonomic data.
- Demonstrated the application and necessity of these preprocessing steps using 1H NMR spectroscopic data from urine.
- Analyzed the impact of "binned" data and normalization methods on correlation coefficient calculations.
Main Results:
- Normalization and scaling are essential preprocessing steps for 1H NMR spectroscopic data from urine.
- "Binned" data present unique challenges in metabonomics analysis.
- Careful consideration is needed when calculating correlation coefficients for normalized data, especially with constant sum normalization.
Conclusions:
- Data preprocessing in metabonomics is highly context-dependent, with no universal optimal method.
- These preprocessing considerations apply to various biofluids, analytical techniques (e.g., HPLC-MS), and other omics fields (e.g., transcriptomics, proteomics).
- Integrated analysis of "fused" datasets requires analogous preprocessing considerations.