Related Experiment Videos
Induction of comprehensible models for gene expression datasets by subgroup discovery methodology
Dragan Gamberger1, Nada Lavrac, Filip Zelezný
1Laboratory for Information Systems, Rudjer Bosković Institute, Zagreb, Croatia. dragan.gamberger@irb.hr <dragan.gamberger@irb.hr>
Journal of Biomedical Informatics
|October 7, 2004
Summary
This study introduces a new method for finding disease markers from gene expression data, creating simple, interpretable logic-based classifiers. This approach enhances robustness and offers novel biological insights, overcoming limitations of complex machine learning models.
Area of Science:
- Bioinformatics
- Machine Learning
- Genomics
Background:
- Machine learning for disease classification from gene expression data faces overfitting due to high dimensionality and limited samples.
- Current complex classifiers offer poor transparency, hindering biological interpretation.
Purpose of the Study:
- To develop simple, robust, and interpretable logic-based classifiers for gene expression data.
- To apply subgroup discovery methodology to gene expression classification problems.
Main Methods:
- Utilized a recently developed subgroup discovery methodology.
- Applied the methodology to two publicly available gene expression datasets.
Main Results:
- Demonstrated the feasibility of constructing simple yet robust logic-based classifiers.
- Discovered classifiers that are amenable to direct expert interpretation.
- Identified classifiers offering novel biological interpretations.
Conclusions:
- Simple logic-based classifiers can be effectively derived from gene expression data.
- This approach enhances transparency and biological insight compared to complex models.
- Subgroup discovery offers a promising methodology for robust disease marker identification.