Related Experiment Videos
Gene expression data analysis of human lymphoma using support vector machines and output coding ensembles
1Dipartimento di Informatica e Scienze dell'Informazione, Università di Genova, via Dodecaneso 35, 16146 Genova, Italy. valenti@disi.unige.it
Artificial Intelligence in Medicine
|November 26, 2002
Summary
This study applies advanced machine learning, including Support Vector Machines (SVM) and Output Coding (OC) ensembles, to DNA microarray data for classifying lymphoma types and identifying distinct disease subgroups.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- DNA microarrays generate vast datasets, initially analyzed with unsupervised methods.
- Supervised methods like decision trees, Support Vector Machines (SVM), and Multi-Layer Perceptrons (MLP) are increasingly used for tissue classification.
Purpose of the Study:
- To classify normal versus tumorous tissues using non-linear SVM and OC ensembles.
- To classify different lymphoma subtypes.
- To analyze coordinately expressed genes in lymphoid tissue carcinogenesis.
Main Methods:
- Utilized non-linear Support Vector Machines (SVM) with polynomial and Gaussian kernels.
- Employed Output Coding (OC) ensembles of learning machines.
- Analyzed gene expression data from the specialized "Lymphochip" DNA microarray.
Main Results:
- SVM successfully distinguished normal from tumorous tissues.
- OC ensembles effectively classified various lymphoma types.
- Identified a gene expression signature differentiating two subgroups within diffuse large B-cell lymphoma (DLBCL).
Conclusions:
- Non-linear SVM and OC ensembles are effective tools for cancer genomics research.
- The findings support the hypothesis of distinct disease entities within DLBCL.
- Gene expression analysis provides insights into lymphoid tissue carcinogenesis.