Related Experiment Video
Updated: Aug 6, 2026

12:13
Sequencing Small Non-coding RNA from Formalin-fixed Tissues and Serum-derived Exosomes from Castration-resistant Prostate Cancer Patients
Published on: November 19, 2019
Machine Learning Classification of Prostate Cancer Genomic Sequences Using K-Mer and Sequence-Derived Features
Kuldeep Rawat1, Hirendra Nath Banerjee2, Jamie Noble2
1Department of Mathematics, Computer Science, and Engineering Technology, Elizabeth City State University, Elizabeth City, USA.
Summary
This study developed a machine learning model to classify cancerous versus healthy prostate cancer (PCa) genomic DNA sequences. The Random Forest model achieved 97.2% accuracy, showing promise for improved PCa diagnostics.
Area of Science:
- Genomics
- Bioinformatics
- Machine Learning
Background:
- Prostate cancer (PCa) disproportionately affects African American men, leading to higher mortality.
- Current diagnostic methods (PSA testing, biopsy) have limitations in specificity and sensitivity.
- Accurate, molecular-level classification tools are needed for early and precise PCa detection.
Purpose of the Study:
- To develop and evaluate a machine learning framework for classifying genomic DNA sequences as cancerous or healthy.
- To identify key genomic features predictive of prostate cancer.
- To assess the potential for equitable, improved diagnostic tools for high-risk populations.
Main Methods:
- A dataset of 1662 high-quality FASTA-formatted DNA sequences from GenBank was analyzed.
- Feature engineering extracted 67 attributes, including GC content, Shannon entropy, and k-mer frequencies.
- A Random Forest classifier was optimized and validated using cross-validation, demonstrating high performance.
Main Results:
- The optimized Random Forest classifier achieved 97.2% cross-validation accuracy and an ROC-AUC of 0.974.
- Key predictive features included sequence length, Shannon entropy, and specific trinucleotide motifs (TTC, AAC, ACC, GGG).
- The model showed high cancer-class recall (0.96) and weighted F1-score (0.95).
Conclusions:
- Interpretable machine learning applied to genomic sequences shows significant potential for prostate cancer classification.
- This approach offers a promising avenue for developing more accurate and equitable diagnostic tools.
- Genomic sequence features can serve as effective biomarkers for PCa detection.
