Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Gene selection and classification of microarray data using random forest.

Ramón Díaz-Uriarte1, Sara Alvarez de Andrés

  • 1Bioinformatics Unit, Biotechnology Programme, Spanish National Cancer Centre (CNIO), Melchor Fernandez Almagro 3, Madrid, 28029, Spain. rdiaz@ligarto.org

BMC Bioinformatics
|January 10, 2006
PubMed
Summary

Random forest classification effectively selects small gene sets for microarray data analysis, maintaining predictive accuracy. This method is suitable for multi-class problems and offers a robust alternative for gene selection in diagnostics.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Evolutionary accumulation modeling in AMR: machine learning to infer and predict evolutionary dynamics of multi-drug resistance.

mBio·2025
Same author

Case of Refractory Classic Hodgkin Lymphoma With a Germline Pathogenic Monoallelic Variant in the TNFRSF13B Gene.

Pediatric blood & cancer·2024
Same author

A hypercubic Mk model framework for capturing reversibility in disease, cancer, and evolutionary accumulation modelling.

Bioinformatics (Oxford, England)·2024
Same author

Cancer Predisposition Syndromes in Children: Who, How, and When Should Genetic Studies Be Considered?

Journal of pediatric hematology/oncology·2024
Same author

HyperTraPS-CT: Inference and prediction for accumulation pathways with flexible data and model structures.

PLoS computational biology·2024
Same author

A variant of the gene <i>HARS</i> detected in the clinical exome: etiology of a peripheral neuropathy undiagnosed for 20 years.

Advances in laboratory medicine·2023

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Genomics

Background:

  • Gene selection is crucial for accurate sample classification in gene expression studies.
  • Traditional methods often use univariate rankings and are limited to two-class problems.
  • Random forest offers advantages for microarray data, handling noise and multi-class issues.

Purpose of the Study:

  • To evaluate random forest for classifying microarray data, including multi-class problems.
  • To introduce a novel gene selection method utilizing random forest.
  • To compare the performance of random forest with existing classification algorithms.

Main Methods:

  • Utilized simulated and nine real-world microarray datasets.
  • Applied random forest for classification and gene selection.

Related Experiment Videos

  • Compared results against Discriminant Analysis (DLDA), K-Nearest Neighbors (KNN), and Support Vector Machines (SVM).
  • Main Results:

    • Random forest demonstrated comparable performance to DLDA, KNN, and SVM.
    • The proposed random forest-based gene selection yielded significantly smaller gene sets.
    • Predictive accuracy was preserved even with reduced gene sets.

    Conclusions:

    • Random forest is a high-performing classification algorithm for microarray data.
    • Gene selection using random forest is effective and produces concise gene lists.
    • Random forest-based methods should be integrated into standard bioinformatics toolkits for class prediction and gene selection.