Related Experiment Video
Updated: May 11, 2026

11:01
Immunoglobulin G N-Glycan Analysis by Ultra-Performance Liquid Chromatography
Published on: January 18, 2020
Greedy feature selection for glycan chromatography data with the generalized Dirichlet distribution
Marie C Galligan1, Radka Saldova, Matthew P Campbell
1School of Mathematical Sciences, University College Dublin, Belfield, Dublin 4, Ireland. mgalliga@tcd.ie
BMC Bioinformatics
|May 9, 2013
Summary
This study introduces a new statistical method for selecting features in compositional glycan data, improving cancer biomarker discovery. The approach enhances accuracy in differentiating between disease states using serum glycan profiles.
Area of Science:
- Biochemistry and Molecular Biology
- Bioinformatics and Computational Biology
- Translational Medicine
Background:
- Glycoproteins and their glycosylation patterns are crucial in biological processes and disease development, especially cancer.
- Serum glycome analysis holds potential for identifying novel cancer biomarkers for early detection.
- High-throughput hydrophilic interaction liquid chromatography (HILIC) generates compositional glycan data requiring specialized statistical analysis.
Purpose of the Study:
- To develop and present a novel methodology for feature selection in compositional glycan data.
- To provide a statistical framework for analyzing glycan chromatography datasets to identify potential cancer biomarkers.
- To address the unique mathematical properties of compositional data in biomarker discovery.
Main Methods:
- A greedy search algorithm utilizing the generalized Dirichlet distribution was employed for feature selection.
- Compositional variables were modeled using beta distributions to identify discriminatory "grouping variables".
- The proposed method was applied to two glycan chromatography datasets and compared against correlation-based feature selection (CFS) and recursive partitioning (rpart).
Main Results:
- The developed feature selection methodology demonstrated effective performance on both glycan chromatography datasets.
- The proposed method achieved a lower misclassification rate compared to CFS and rpart.
- Higher sensitivity rates were observed with the proposed method, indicating improved biomarker identification capabilities.
Conclusions:
- The novel feature selection method is effective for analyzing compositional glycan data.
- While computationally more intensive, the proposed method offers superior classification accuracy and sensitivity for biomarker discovery.
- This approach provides a valuable tool for identifying glycan biomarkers in complex biological samples.

