Related Experiment Videos
Joint analysis of two microarray gene-expression data sets to select lung adenocarcinoma marker genes
Hongying Jiang1, Youping Deng, Huann-Sheng Chen
1Department of Mathematical Sciences, Michigan Technological University, Houghton, MI 49931, USA. hojiang@mtu.edu
BMC Bioinformatics
|June 26, 2004
Summary
This study integrated two lung cancer microarray datasets to identify diagnostic and survival marker genes. The developed methods accurately predict patient outcomes and aid in cancer diagnosis.
Area of Science:
- Bioinformatics
- Genomics
- Cancer Research
Background:
- Microarray experiments face challenges with high costs, low reproducibility, and limited sample sizes.
- Integrating multiple microarray datasets increases statistical power and aids in identifying common marker genes for diseases like lung cancer.
Purpose of the Study:
- To develop and validate a method for integrating disparate microarray datasets.
- To identify novel marker genes for distinguishing lung cancer from normal samples.
- To identify genes significantly correlated with patient survival in lung cancer.
Main Methods:
- Combined two lung cancer GeneChip microarray datasets using a distribution transformation method.
- Applied gene shaving (GS) methods based on Random Forests (RF) and Fisher's Linear Discrimination (FLD) for marker gene selection.
- Utilized a two-step survival test, including univariate Cox proportional hazard regression and principal component regression, to identify survival-related genes.
Main Results:
- Identified 13 (RF) and 10 (FLD) marker genes differentiating lung cancer from normal samples, with 5 common genes.
- Classifiers built from one dataset predicted the other with >98% accuracy.
- 36 genes showed significant correlation with patient survival, with 26 previously reported, including tumor suppressors and oncogenes. Principal component regression reduced this to 16 genes.
Conclusions:
- Developed a robust method for integrating microarray data from different sources.
- Demonstrated the effectiveness of gene shaving and survival analysis for identifying minimal marker gene sets for cancer diagnosis and prognosis.
- High prediction accuracy achieved when applying classification models developed from one integrated dataset to another.