Related Experiment Video
Updated: Jun 11, 2025

07:41
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
8.9K
Evolutionary Mechanism Based Conserved Gene Expression Biclustering Module Analysis for Breast Cancer Genomics
Wei Yuan1, Yaming Li1, Zhengpan Han1
1School of Biomedical Engineering, Guangzhou Medical University, Guangzhou 511436, China.
Biomedicines
|September 28, 2024
Summary
A new algorithm, Conserved Gene Expression Module based on Genetic Algorithm (CGEMGA), efficiently identifies significant gene biclusters and functionally related genes from RNA-sequencing data, outperforming existing methods in speed and accuracy for cancer research.
Area of Science:
- Bioinformatics and Computational Biology
- Genomics and Transcriptomics
- Cancer Research
Background:
- RNA sequencing generates vast gene expression data, necessitating methods to identify significant gene biclusters and functionally related genes.
- Existing algorithms for gene expression module discovery face challenges in efficiency and accuracy when analyzing large datasets.
Purpose of the Study:
- To propose a novel algorithm, Conserved Gene Expression Module based on Genetic Algorithm (CGEMGA), for identifying significant gene biclusters.
- To evaluate the performance of CGEMGA against existing methods using breast cancer data from The Cancer Genome Atlas (TCGA) database.
- To assess the biological relevance of identified gene modules through enrichment analysis of driver genes and cancer-related pathways.
Main Methods:
- Development and implementation of the Conserved Gene Expression Module based on Genetic Algorithm (CGEMGA).
- Application of CGEMGA to breast cancer RNA-sequencing data from the TCGA database.
- Statistical evaluation using Fisher's exact test (p-values) and F-test to compare CGEMGA with other algorithms (e.g., Cheng and Church, CGEM).
- Analysis of computational cost by measuring algorithm running times.
- Enrichment analysis to identify driver genes and cancer-related pathways.
Main Results:
- CGEMGA achieved a superior average p-value of 1.54 × 10-4 ± 3.06 × 10-5 across 10 independent runs, indicating higher significance.
- The F-test showed a significant difference between CGEMGA and the CGEM algorithm (p-value = 0.039).
- CGEMGA demonstrated significantly shorter average runtime (5.22 × 100 ± 1.65 × 10-1 s), highlighting its computational efficiency.
- Enrichment analysis confirmed that genes identified by CGEMGA are significantly enriched for driver genes.
Conclusions:
- CGEMGA is a fast, robust, and efficient algorithm for extracting co-expressed genes and associated biclusters from RNA-seq data.
- The algorithm demonstrates superior performance in identifying significant gene modules compared to existing methods.
- CGEMGA provides a valuable tool for cancer research by effectively identifying biologically relevant gene sets.

