Related Experiment Video
Updated: Sep 9, 2025

Microarray-based Identification of Individual HERV Loci Expression: Application to Biomarker Discovery in Prostate Cancer
Published on: November 2, 2013
Addressing wide-data studies of gene expression microarrays with the Relevance Feature and Vector Machine
Albert Belenguer-Llorens1, Carlos Sevilla-Salcedo1, Emilio Parrado-Hernández1
1Universidad Carlos III de Madrid, (1), Department of Signal Processing and Communications, Leganés, 28911, Spain.
None:
This paper presents the Relevance Feature and Vector Machine (RFVM), a novel Bayesian model that addresses the wide-data challenges of gene expression microarrays, ensuring interpretability in both feature and sample spaces. The wide-data problem occurs when Machine Learning algorithms encounter databases with significantly more features than observations, as commonly seen in gene expression microarrays. To address these challenges, RFVM operates in the dual space, enabling effective parameter inference with small patient cohorts, thereby avoiding overfitting and ensuring generalizable solutions. The core innovation of RFVM lies in its two-way sparsity approach, incorporating priors over both primal and dual variables to perform joint feature and sample selection. As we will show, this capability is critical in wide-data clinical settings, as identifying relevant patients enhances the feature selection process, while effective feature selection, in turn, improves sample identification. This interaction results in more compact solutions and boosts the interpretability and performance of the final model. The RFVM's capabilities are validated against several models on multiple gene expression microarray datasets. Results demonstrate that RFVM achieves superior performance in diagnostic tasks while producing the most compact solutions. Furthermore, the selected genes align with known biomarkers from the medical literature, highlighting the potential of the model as a clinical tool.

